CrawlCheck

Findings · 2026-09-30 · By · 0 views

Our extension read a page as 36.5% JavaScript. It was zero.

How a comparison of two copies of the same 2,850 words came out a third wrong: undecoded entities and a signed-in admin bar. What version 1.0.2 compares instead, and the controls that prove it.

On 25 September 2026 the CrawlCheck browser extension reported 36.5 percent of a homepage as JavaScript-dependent; the true figure was zero, 2,850 words identical in raw and rendered copies. Two causes: typographic entities such as ’ were undecoded in the raw copy so every sentence with an apostrophe mismatched (36.5 to 11.9), and the signed-in owner's WordPress admin bar was counted as page text (11.9 to 0). Version 1.0.2 decodes every entity, compares word multisets and strips the admin bar; controls read 49.9 percent and 100 percent as expected.

Share of all scans carrying each finding named aboveCONTENT_NEEDS_JAVASCRIPT0.6%Share of all scans carrying eachfinding named aboveCONTENT_NEEDS_JAVASCRIPT0.6%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

The CrawlCheck browser extension has one measurement no server-side scanner can make: it reads the page from inside a real browser, after JavaScript has run, and compares that to the raw HTML the server sent. The difference is the share of the page's text that only exists after rendering, which is the share a crawler that does not execute JavaScript never sees. On 25 September the extension reported that share as 36.5 percent for the homepage of a site we run ourselves. The true figure was zero: 2,850 words in the raw HTML, 2,850 words in the rendered page, identical. This is how a comparison of two copies of the same text came out a third wrong, in two separate ways, and what the fixed version does instead.

The measurement #

The extension takes the HTML the server delivered, extracts its visible text, and does the same to the live document after the page has finished loading. It then asks which sentences in the rendered copy are absent from the raw copy. Sentences present only after rendering are JavaScript-dependent; their share of the whole is the number on the popup. The method is sound. Both of its inputs were wrong, in ways that did not cancel.

Defect one: two spellings of an apostrophe #

Raw HTML on a WordPress site carries typographic punctuation as entities: ’ for a right single quotation mark, — for a dash, → for an arrow. The live DOM, read through the browser, has already decoded them into the characters themselves. The extension's text extractor decoded some entities and not these, so every sentence containing an apostrophe, a dash or an arrow existed in two spellings: one with ’ in the raw copy and one with ’ in the rendered copy. An exact sentence match called every one of them a sentence that only existed after rendering. On a page written in ordinary English prose that is most of the sentences. Decoding every entity before comparing took the figure from 36.5 percent to 11.9.

Defect two: the person running the test was signed in #

The remaining 11.9 percent was real text that was in the rendered page and not in the raw HTML, and it was real because WordPress adds an administration bar to the top of every page for a signed-in site owner. That bar is injected by script, it carries a few hundred words of menu labels, and it is never served to anyone who is not signed in, crawlers included. The extension counted it as JavaScript-dependent content, which it is, on a page that only the owner ever sees that way. Stripping the admin bar from the rendered snapshot took the figure from 11.9 percent to zero, which is the correct answer for that page.

StepReported JS-only shareCause
as shipped in 1.0.136.5%entities undecoded in the raw copy; exact sentence match
after full entity decode11.9%signed-in admin bar counted as page text
after admin bar stripped0%2,850 of 2,850 words identical
control: injected paragraph49.9%reads as 'most' — correct
control: empty shell, all text by script100%correct

What the fixed version compares #

Version 1.0.2 does three things differently. Every entity is decoded in both copies before any comparison. The comparison is no longer an exact sentence match but a difference of word multisets: the words of the rendered copy minus the words of the raw copy, counted with multiplicity, so a sentence that differs only in punctuation contributes nothing and a sentence that was inserted contributes its words. And the administration bar is removed from the rendered snapshot before extraction. Two controls were run against the fix, because a fix that makes one number go to zero has to be shown not to make every number go to zero: a page with a paragraph injected by script reads 49.9 percent, and a page whose entire body is written by script reads 100. Both are the right answers.

Why the two defects did not cancel #

A measurement error that only ever inflates is easier to catch than one that could go either way, and both of these inflated. The entity mismatch could only ever move sentences from the shared column into the rendered-only column, never the reverse, because the raw copy is the one with the entities in it. The admin bar could only ever add words to the rendered copy. So a page with no JavaScript-dependent text at all could not read zero under 1.0.1 if it had an apostrophe in it and its owner was signed in, which describes nearly every WordPress homepage measured by its own owner. The reading was biased upward for the exact population most likely to run the tool.

Why the first number looked plausible #

A third of a page depending on JavaScript is a common enough finding on modern sites that 36.5 percent did not read as an error. It read as a finding. That is the failure mode to watch for in any measurement tool: a wrong number that lands inside the range of right numbers is not caught by looking at it. It was caught because the site being measured was one whose raw HTML we know, and a site whose homepage is server-rendered prose cannot be a third client-rendered. The general rule, applied across the image count that was wrong and the placeholder check that accused a newspaper, is that a new reading is verified on a page whose true value is known before it is believed on a page whose value is not.

What it means for the reading you see #

The rendering share shown by the extension from 1.0.2 onward is a word-level difference between two entity-decoded copies with the owner's admin bar excluded. It is still a measurement from one browser at one moment, and a page that renders differently by viewport or by session will read differently. The server-side counterpart, which cannot execute JavaScript at all and reports CONTENT_NEEDS_JAVASCRIPT when the delivered HTML is thin, is described in the rendering guide. The extension's own page and its privacy disclosures are at /extension; the update was accepted by the Chrome Web Store on 29 September.

Live, as you read this: the corpus now holds 7,808 domains across 2,465 scans. The figures in this piece were measured on the date above; this line is not.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What does the extension's JavaScript-only share measure?

The share of a page's visible text that exists in the rendered document but not in the raw HTML the server sent, which is the share a crawler that does not execute JavaScript never sees.

Why did it report 36.5 percent on a page where the true figure was zero?

Two defects: typographic entities such as ’ were left undecoded in the raw copy while the browser had decoded them, so every sentence with an apostrophe mismatched; and the signed-in owner's WordPress admin bar, injected by script, was counted as page text.

How was the correct figure established?

By counting: 2,850 words in the raw HTML and 2,850 in the rendered page, identical after entity decoding and admin-bar removal. Two controls confirmed the fix did not zero everything: an injected paragraph reads 49.9 percent and a script-only page reads 100.

What changed in version 1.0.2?

Full entity decoding of both copies, a word-multiset difference instead of an exact sentence match, and removal of the admin bar from the rendered snapshot. The update was accepted by the Chrome Web Store on 29 September 2026.

Does the reading still depend on who runs it?

Less than before, since the admin bar is excluded, but it is still one browser at one moment. A page that renders differently by viewport or session will read differently.

Why was the wrong number not obvious?

Because a third of a page depending on JavaScript is within the normal range for modern sites. It was caught only because the measured page's raw HTML was known to be complete.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works