CrawlCheck

Findings · 2026-08-13 · By · 0 views

What answer engines receive from local business sites: mostly not text

Three findings from a corpus of sites measured byte by byte, and why none of them appear in a validator.

Mostly not text. When one page is fetched by fifteen client identities from a single address inside one second, what comes back is dominated by markup, script and inline styling — the visible words are a small fraction of the bytes. Some responses return HTTP 200 and still carry nothing a crawler can use, and some origins refuse answer-engine identities at the edge while serving an ordinary browser normally. Every part of the measurement is a plain fetch, which means you can repeat it on your own site and disagree with it.

How much of a delivered page is text a machine can quote
A page we measured0.8% of 2.3MB
A typical local site18.8%
What we look for20% and up

Bars are drawn to scale. The rest of every bar is code the crawler downloads and cannot use.

Every SEO tool measures what a page says. Almost none measure what a machine is handed when it asks for that page. Those are different numbers, and the gap between them is where most of the damage lives.

How the share of text is measured #

The measurement is deliberately plain. Take the bytes that arrive on the wire for one URL. Strip every script, style, noscript, template and inline svg block, drop the comments, drop the tags, collapse the whitespace. What remains is the text a client that does not run a browser could read — and, if it wanted to, quote. Divide by the bytes it cost. That ratio is stored for every page the depth crawl visits, and the site-level number reported is the median across them, so one heavy template and one light one do not cancel out into a fiction.

Two thresholds sit on that ratio. The content row of the report passes at 20%. Below 5% a separate finding fires, PAGE_IS_MOSTLY_CODE, at a weight that caps the grade — because before it existed a page at 2% and a page at 19% scored identically on the row, and they are not the same page. Across the corpus behind these figures the finding fires on roughly one scan in seven.

Live, as you read this: the corpus now holds 7,768 domains across 2,400 scans. The figures in this piece were measured on the date above; this line is not.

1. The heaviest page we measured delivered 2.3% text #

476,540 bytes on the wire. Visible, readable text: 2.3% of it. The rest was markup, inline styles and scripts. A crawler pays that cost on every URL it fetches, and inline CSS cannot be cached between pages, so it pays again on the next one.

The page looked fine to every human who visited it. It renders quickly, it reads well, and no validator has an opinion about it.

2. A file returning HTTP 200 that no crawler could read #

On one site, /robots.txt answered every request with a browser-verification interstitial and a 200 OK. Crawlers do not execute the JavaScript that clears such a page — they read the interstitial as the file. Every uptime monitor called the site healthy.

The cause was not the security product. It was a cache rule storing whatever came back, including a challenge page, and serving it to everyone for the next five minutes. A per-address defence became a site-wide outage of the machine layer. That one has its own write-up, including the two-request test that catches it without needing to recognise the page.

3. Two engines refused at the edge while Google walked in #

On another site, GPTBot and ClaudeBot both received 502 while Googlebot and PerplexityBot were served normally. Nothing in robots.txt said so — this was the edge, and almost nobody configures it deliberately.

Until that is lifted, nothing else about the site reaches those two engines, however good it is. The scan records this as ANSWER_ENGINE_REFUSED, and it has two quieter siblings that only a side-by-side fetch can see: CRAWLER_SERVED_LESS, when a named crawler is handed materially less text than an ordinary client gets for the same URL, and CRAWLER_REDIRECTED_AWAY, when it is sent somewhere the browser was not. Whatever sits at that destination is what the engine reads and quotes instead of your page.

How much of a delivered page is text a machine can quote #

PageReadable text
A page we measured0.8% of 2.3 MB
A typical local site18.8%
What we look for20% and up

The rest is markup, script and style: bytes that arrive, are paid for, and say nothing an answer engine can quote.

What moved the number on four of our own sites #

We run this scanner against properties we operate, and three of them sat under the 5% line. The fixes were mechanical and the results were measured, not estimated.

On the first, a page builder had emitted the same icon graphics inline dozens of times. Folding the duplicates into one symbol table and stripping comments and inter-tag whitespace took the homepage from 221,487 to 204,666 bytes with the text byte-identical, and the share from 4.78% to 5.17%. On the second, ten inline style blocks were the weight; writing each one over 2 KB out to a cached file and replacing it in place, in the same order so the cascade did not change, took 106,315 bytes to 73,133 and the share from 4.03% to 5.86%. Both crossed the line. Both are now uncapped.

The third did not, and it is the more useful case. That homepage carries 224,485 bytes and 5,485 characters of text — 2.44%. Its ten style blocks hold 73,881 bytes. Strip every one and it lands at 3.61%. Externalise all 31,000 bytes of non-schema script on top of that and it reaches about 4.5%. Still under. Compression cannot fix a page that has nothing to say; it needs words, not fewer bytes. The scanner cannot tell those two situations apart from the ratio alone, which is why the report shows the byte count and the character count side by side rather than only the percentage.

The guard that nearly shipped a silent no-op #

The externaliser first refused to return any output smaller than 55% of its input, as a safety net against a transform going wrong. On the site’s 404 template that is a legitimate 48% reduction, so the guard silently did nothing there and the page kept its weight. A percentage floor is the wrong guard for a transform whose output size is computable: it now tracks what each replacement removes and adds and requires the final length to equal input minus that delta, exactly. The general lesson is the same one the viewport defect taught — a check that proves your code ran tells you nothing about whether it did the right thing.

What this means if you own a site #

None of these three show up in a rank tracker, a validator, or an uptime check. They show up when you fetch the site the way an answer engine fetches it, as the crawlers an answer engine sends, and compare that against what a browser gets. The order of repair follows from the numbers: a refusal at the edge comes first, because nothing behind it is reachable; a challenge page cached as a file comes second, for the same reason; the byte ratio comes last, because a page that is reachable and thin is still reachable.

That comparison is free here, and no scanned domain is ever named in anything we publish — including this article. The terms used above are defined in the glossary.

Fifteen identities, one second, one address #

The measurements above come from a specific method, and the method is the reason the numbers mean anything. The homepage is fetched fifteen times in the same second: twelve named crawler identities, an unnamed client as a control, and two identical mobile-browser requests.

The second browser request is the part that carries the weight. Before any difference between a crawler and a browser can be called a difference, the site has to prove it answers the same question the same way twice. If those two identical requests disagree, the site is non-deterministic — A/B testing, a rotating cache, a personalisation layer — and every comparison downstream is noise. That check runs first, and when it fails the divergence findings are withheld rather than reported.

One address is a limitation, not a feature, and it is stated on every refusal this scanner reports. A site that refuses us may be refusing crawlers by name, or refusing our datacentre range, or having a bad minute. Those are three different facts and a single vantage cannot separate them, which is why a refusal is corroborated against a second network before it is treated as a pattern.

What the bytes actually break down into #

“Text percentage” is a summary of a fuller measurement. Every delivered homepage is decomposed into the categories below, and the interesting number is usually not the text share but which category ate the rest.

CategoryWhat it isCan an engine quote it?
Readable textThe words a reader seesYes — this is the entire quotable surface
Markup and attributesTags, classes, data attributesNo
Structured dataJSON-LD blocksNot as prose, but it is how facts get attached to an entity
Inline CSSStyles in the document rather than a fileNo
Inline JavaScriptScripts in the documentNo
Inline SVGVector graphics written into the HTMLNo, and it is frequently the largest single block
HTML commentsNotes, build stamps, commented-out sectionsNo, and they ship to every client

Inline SVG is the category that surprises people. An icon set written into the document rather than referenced as a file can outweigh the article it sits beside, and it compresses well enough that nobody notices in a page-weight audit. PAGE_IS_MOSTLY_CODE fires on 22.4% of scans, and the breakdown is what turns that finding from a scolding into a task.

The half a server cannot see #

Everything above describes what arrives before any JavaScript runs, because that is what a crawler fetching your URL receives. It is not the whole story, and saying so matters more than the number.

Content that only exists after rendering is a separate question with a separate code, CONTENT_NEEDS_JAVASCRIPT. A page can be 3% text on the wire and perfectly readable in a browser, and a page can be text-rich on the wire and still fail for a crawler that is refused at the edge. These are independent failures. Treating a low text share as proof of invisibility, or a high one as proof of health, is the mistake this measurement is most often used to make.

Two vantages settle what one cannot. The scan reads what an anonymous non-browser client is served from a datacentre address; the browser extension reads what is served to your own address, after your JavaScript has run. Where they disagree is where the interesting answer lives, and neither one alone can find it.

Reading this on your own site without any tooling #

You can reproduce the core of it with two commands and no account. Fetch your homepage with a plain HTTP client and save the body. Fetch it again with a browser user-agent string and save that. Compare the byte counts. If they differ materially, something in your stack is making a decision based on who asked — which may be a CDN rule you inherited, a bot-management product, or a cache that segments on user-agent.

Then strip the tags from the first file and count what is left. That number, over the total, is the share of what you delivered that a machine can actually quote. It is a blunt instrument compared to the breakdown above, and it will still tell you within a minute whether you have a problem worth the afternoon.

Why the control request is the whole method #

It is worth being explicit about why a control matters, because most tools that compare crawler and browser output do not use one, and their divergence numbers are therefore unfalsifiable.

Suppose a site serves 90 KB to a browser and 40 KB to GPTBot. That looks like cloaking. Now suppose the same site serves 90 KB to one browser request and 45 KB to an identical browser request sent in the same second — which happens on any site running a split test, a rotating banner, or a cache that has not warmed. The crawler gap is now inside the site’s own variance, and calling it discrimination would be a false accusation dressed in a number.

The control makes the difference checkable. Two identical requests establish the noise floor; any comparison against a named crawler is only reported when it clears that floor. It costs one extra fetch and it is the difference between a measurement and an opinion.

The same discipline applies to refusals. A request that is turned away tells you that this address, at this moment, using this name, was refused. It does not tell you that crawlers are blocked. Those are different claims, and only the first one is supported by a single vantage — which is why a refusal measured once is reported as an observation and only becomes a finding when it repeats.

How to see what a crawler receives from your own site #

Fetch your homepage three times: as a browser, as GPTBot, and with no user-agent at all. Compare status, byte count and content type. Then strip the markup and count what is left:

for ua in 'Mozilla/5.0' 'GPTBot' ''; do curl -s -A "$ua" -o /tmp/p -w "$ua %{http_code} %{size_download}\n" https://yoursite.com/; done
sed 's/<[^>]*>/ /g' /tmp/p | wc -w

Three different byte counts mean three different pages. A word count that is a small fraction of the byte count is the 2.3% problem above. How to check if AI crawlers can read your site walks through the full version; do AI crawlers render JavaScript explains why the served document is the only one that counts.

Where each of the three appears in a report #

The payload section scores the text-to-bytes ratio and the render path. The machine-layer and trust-chain sections read robots.txt by content type, not status, after the 200 that could not be read. The agent-view section fetches the homepage as fifteen named identities and reports which were refused, challenged, or handed less text than an unnamed client.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

How much of a web page is text a machine can quote?

On the heaviest page measured, 2.3%. The rest was markup, script and styling that an answer engine cannot cite.

Why do validators miss this?

A validator checks whether markup is well formed. It does not weigh how much of the delivered payload is quotable text.

What should a site owner do about it?

Measure the ratio on your own heaviest pages before adding anything else. Weight you remove is quotable text you gain, proportionally.

How much of a web page is actually text?

Far less than most people assume. On the local-business pages measured here the visible words are a small share of the delivered bytes, with the rest going to markup, inline styling and script. The practical test is to fetch the page, strip the tags and compare what is left against the byte count, which is the same ratio the scanner records.

Do AI crawlers get the same page a browser gets?

Usually, but not always, and the exceptions are invisible from a browser. The way to find out is to request the same URL under several client identities at the same moment from the same address and compare the responses byte for byte. Identical bytes everywhere means uniform delivery. Divergence means something in your stack is deciding what to serve based on who asked.

Why does a page return 200 and still give a crawler nothing?

Because a status code describes the transaction, not the content. A 200 can carry a challenge page, an empty shell awaiting JavaScript, a soft error page, or a machine file served as HTML. All four look like success in a monitoring dashboard and deliver nothing readable, which is why the status has to be read alongside what the body actually contains.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works