CrawlCheck

Guides · 2026-10-02 · By · 0 views

Text-to-HTML ratio: when a page is mostly code

Text-to-HTML ratio is not a ranking factor, but a page whose visible words are under five percent of its bytes is a page where crawlers download framework state, inline CSS and SVG before they reach a sentence. How to measure it, what the threshold means, and what to remove first.

Text-to-HTML ratio is the share of a page's decompressed HTML bytes that is visible text. Google has said it is not a ranking signal, but the number still matters for machines with byte and token budgets: a crawler or answer engine fetching a page that is 3 percent text spends 97 percent of the transfer on inline CSS, scripts, JSON state and SVG. CrawlCheck raises PAGE_IS_MOSTLY_CODE on large pages whose visible text is a small share of the bytes, and names the largest single block so the fix starts there.

Share of all scans carrying each finding named abovePAGE_IS_MOSTLY_CODE29.2%Share of all scans carrying eachfinding named abovePAGE_IS_MOSTLY_CODE29.2%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Text-to-HTML ratio has a bad reputation in SEO, deservedly: Google has said repeatedly that it does not use the ratio as a ranking signal, and tools that turned it into a score with a target percentage invented a rule. But the measurement underneath is useful for a different reason. Every machine that reads a page pays for bytes and, in the case of language models, for tokens. A page whose visible words are a few percent of its HTML makes every reader download a framework before it reaches a sentence. This guide covers how to measure the ratio honestly, what the threshold means, and what to remove first.

Where the bytes of a homepage go, and the 5 percent lineDECOMPRESSED HTML OF ONE HOMEPAGE, 100% = ALL BYTESinline CSS 31%inline JSON / state 24%inline scripts 21%SVG markup 9%attributes, tags 11%visible text 4%WHAT THE FINDING NEEDSvisible text < 5%of decompressed bytespage > 50,000 bytesdecompressedPAGE_IS_MOSTLY_CODEnames the largest single blockIllustrative split. A small page is not judged on ratio; a large one with under 5% text is.

What the ratio measures #

Take the HTML exactly as the server sends it, decompress it, and count the bytes. Then extract the text a person would see, excluding scripts, styles, templates and hidden elements, and count those bytes. The ratio is the second over the first. Measuring compressed bytes instead would flatter repetitive markup, because CSS and JSON compress very well; the decompressed size is what a parser actually walks and what a tokenizer consumes.

The ratio alone is not a verdict. A 6 KB page that is 3 percent text is a short page with a nav bar, and there is nothing to fix. A 400 KB page that is 3 percent text is a page carrying a serialized application state, an inlined stylesheet and an icon set before its first paragraph. Size and share have to be read together.

The threshold, and why it is not a target #

Under 5 percent on a large page is plainly unusual. That line is a floor, not a goal; a page at 6 percent is not healthy because it cleared it. CrawlCheck raises PAGE_IS_MOSTLY_CODE only on large pages with a low text share, never on small pages where a ratio means nothing. Across the live corpus the finding fires on 29.2% of scans, which makes it one of the most common things the scanner reports.

Largest blockWhere it usually comes fromFirst fix
inline <style>page builders and critical-CSS plugins that inline everythinginline only above-the-fold rules; move the rest to a cached file
inline JSON stateframework hydration data (__NEXT_DATA__, Nuxt, Gatsby page data)trim the props the page does not render; paginate large lists
inline <script>tag managers, chat widgets, consent tools pasted into the headload from a file with defer; drop unused tags
SVG markupicon sets inlined on every pageone sprite file referenced by <use>, or image files
data URIsimages and fonts base64-encoded into CSS or HTMLserve as files so they cache and do not inflate the page
repeated markupmega-menus and footers with hundreds of linksshorten the menu; render rarely used parts on demand

Why it matters for answer engines in particular #

Search crawlers fetch large pages routinely and have done so for years. Answer engines and agents work differently: they often fetch a page at the moment a user asks, under a time limit and a size limit, and many truncate the response before parsing. A page that front-loads 200 KB of styles and state can be cut off before the text that answers the question. The same page also costs more to convert into model input. None of this is a ranking factor in the classic sense; it is whether the words arrive at all. If the words exist only after JavaScript runs, the problem is different and larger, covered in whether AI crawlers render JavaScript.

Measured on a site we run #

A tree-service homepage we operate measured 7.7 percent visible text before a payload pass, with seven render-blocking stylesheets and the page builder's inline bootstrap script in the head. Merging four adjacent per-section stylesheets into one cached file, deferring the script loader, moving an inline compatibility script to a file and removing an emoji script, together with about 450 words added to an on-page guide, took it to 10.3 percent. Most of the movement came from deleting bytes that were not content, which is the general pattern: the text is rarely too short, the wrapper is too large.

How to fix it without breaking the page #

Start with the largest single block, because the distribution is almost always lopsided: one inlined stylesheet or one hydration payload is often half the page. Move inline CSS to an external file the browser caches, keeping only the rules needed to paint the first screen. Cut hydration data down to what the page renders. Replace inlined icon SVGs with a sprite. Then re-measure the decompressed size and the text share, and check that the page renders identically. Our caching guide covers the edge-cache side, which decides whether a slimmer page is actually what crawlers receive, and the AEO audit page reads this finding next to the other delivery checks.

Measuring it yourself #

You can reproduce the measurement with two commands. Fetch the page with a plain client that accepts compression, decompress it, and count the bytes; then strip script, style, noscript, template and svg elements, remove the remaining tags, collapse whitespace and count again. The ratio is the second count over the first. Do it on the HTML the server sends, not on the browser's live DOM, because the live DOM includes text inserted by JavaScript that a non-rendering crawler never receives, and excludes the inline state blobs that make up much of the transfer. Run it on the homepage and on your heaviest template, which is often a product listing or a location page rather than the homepage.

Then list the largest single elements by byte size. In our scans the largest block is usually far larger than the second: an inline stylesheet of several hundred kilobytes, a hydration script carrying the data for every product in a category, or a mega-menu with hundreds of links repeated in a mobile copy. Fixing that one block typically moves the ratio more than every other change combined. Re-measure after each change, and re-measure through the CDN, because a cached copy of the old page will keep serving the old ratio until it is purged.

What the measurement does not say #

A high ratio does not mean the page is good, only that it is not mostly wrapper. A parked domain with 800 words of boilerplate has an excellent ratio, which we covered in a parking page that defeats every thin-content check. The ratio is a delivery measurement, not a quality one.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Is text-to-HTML ratio a Google ranking factor?

No. Google has said it does not use the ratio as a ranking signal. It is still a useful measure of how much of a page's transfer is content rather than code.

What is a good text-to-HTML ratio?

There is no target. Visible text under 5 percent on a large page is unusual enough to be worth fixing.

Should the ratio use compressed or decompressed bytes?

Decompressed. CSS and JSON compress very well, so compressed sizes understate how much markup a parser or tokenizer has to process.

What usually makes a page mostly code?

Inlined stylesheets, framework hydration state, tag-manager and widget scripts in the head, inlined SVG icon sets and base64 data URIs. One of them is usually most of the page.

Why does it matter for AI crawlers?

Answer engines often fetch a page when a user asks, under time and size limits, and may truncate it. A page that front-loads hundreds of kilobytes of code can lose the text before it is read.

Does a high ratio mean the page is good?

No. A parked page full of boilerplate text can have a high ratio. The ratio measures delivery, not quality.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works