CrawlCheck

Findings · 2026-10-06 · By · 0 views

139,788 homepages in four days: what AI crawlers get

The first four days of the open dataset's new lanes, de-duplicated to one reading per site: how many sites block AI crawlers, how many publish llms.txt, and how little of a homepage is text.

Across 139,788 homepages measured between 3 and 6 October 2026, 6.3% block at least one named AI crawler in robots.txt and 15.4% publish an llms.txt file. The median homepage is 3.6% visible text by bytes, 61.5% are under 5%, 28.2% have no sitemap a crawler can find, and 48.4% send an HSTS header. Government and education sites trail on llms.txt, sitemaps, speed and grades.

Between 3 and 6 October 2026 the CrawlCheck dataset read 208,295 distinct domains across four lanes: SaaS and software, government and education, US local businesses, and everything else the crawl reached. 139,788 of them answered well enough to be measured end to end; the rest refused, timed out or never served a homepage, and they are left out of every figure below rather than counted as failures. Where a site was read more than once, only its latest reading counts. Nothing here is estimated and no site is named.

The September edition of this report worked from 1,364 deep scans. This one works from a population a hundred times larger, measured more lightly, so the questions are narrower: what a crawler is allowed to fetch, what it is offered, and what it gets when it arrives.

The four lanes side by side #

MeasureSaaS and softwareGovernment and educationUS local businessesEverything elseAll
Homepages measured19,9829,98420,04289,780139,788
Block at least one AI crawler4.0%2.1%2.3%8.1%6.3%
Publish llms.txt28.2%5.5%22.6%12.0%15.4%
Send an HSTS header54.8%46.1%43.3%48.3%48.4%
No sitemap a crawler can find27.2%33.5%13.5%31.1%28.2%
Homepage under 5% visible text60.0%58.5%72.6%59.7%61.5%
Median time to first byte221 ms418 ms209 ms403 ms333 ms
Grade A25.5%9.9%27.2%15.9%18.5%

Finding 1: blocking AI crawlers is rare #

For all the attention it gets, an explicit AI-crawler block is uncommon. 6.3% of measured homepages name at least one AI crawler in robots.txt and disallow it. GPTBot is the most blocked at 3.7%, followed by CCBot (3.1%), Bytespider (3.0%) and ClaudeBot (2.7%). The detail by crawler, including the split between training and search agents, is in the companion post.

The lanes differ more than the headline suggests. Government and local-business sites block least (2.1% and 2.3%); the broad everything-else lane, which includes publishers and media, blocks most at 8.1%. A small business that has never touched its robots.txt is almost certainly open to every AI crawler, whether it meant to be or not.

Finding 2: llms.txt is common where builders add it #

15.4% of measured sites publish an llms.txt file. SaaS sites lead at 28.2%, and US local businesses are close behind at 22.6%, which is not because small businesses are early adopters: much of that share arrives through website builders that publish the file by default. Government and education sites trail at 5.5%. A file that exists by default says nothing about whether anyone wrote it, so presence is a weak signal on its own. What matters is whether the links in it resolve, which the testing guide walks through alongside robots.txt.

Finding 3: most homepages are almost entirely code #

The median homepage is 3.6% visible text by bytes. 61.5% are under 5%, 30.3% are under 2%, and only 15.6% reach 10%. A crawler pays for every byte it downloads and can only quote the text, so a homepage that is 97% markup, styles and scripts gives an answer engine very little to work with for the size of the download.

Client-side rendering makes it worse. 17.1% of homepages build their content in the browser, and their median text share is 0.3%, against 4.6% for pages rendered on the server. A crawler that does not run scripts sees almost nothing on a client-rendered page. Local-business sites are the heaviest: 72.6% are under 5% text and the median homepage weighs 151 KB. The text-to-HTML guide covers what usually causes this and how to cut it.

Finding 4: a quarter of sites give crawlers no map #

28.2% of measured sites have no sitemap a crawler can find, neither declared in robots.txt nor at the standard locations, and 29.9% do not declare one in robots.txt at all. Government and education sites are worst at 33.5%. Local businesses do best at 13.5%, most likely because the same builders that add llms.txt also generate a sitemap. 12.9% serve the same homepage on two hosts without redirecting one to the other, which splits whatever a crawler learns about the site between two addresses; on government sites it is 23%. The sitemap audit guide covers both.

Finding 5: half the web still has no HSTS #

95.6% of measured homepages answer over HTTPS, but only 48.4% send a Strict-Transport-Security header, so on the rest a visitor's first request can still be downgraded. It is one header line. The HSTS guide explains how to add it safely.

Speed and weight #

The median time to first byte is 333 ms, and 21.5% of sites take longer than a second. Government sites and the everything-else lane are about twice as slow as SaaS and local-business sites (418 ms and 403 ms against 221 ms and 209 ms). The median homepage is 120 KB, and 11.7% are over 500 KB.

One figure left out on purpose #

This dataset also counts robots.txt files whose named groups drop rules the wildcard group carries. That count is not published here: while preparing this report we found that our check overcounted sites that block a crawler outright, and fixed it the same day. The correction explains what was wrong. The corrected rate will appear once the affected sites have been re-read.

Check your own site #

Every figure above is a share of sites. Your own site is either in the majority or it is not, and a scan answers that in under a minute. The daily files behind this report are described on the datasets page.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

How many websites block AI crawlers?

In the CrawlCheck dataset for 3-6 October 2026, 6.3% of 139,788 measured homepages block at least one named AI crawler in robots.txt. GPTBot is the most blocked at 3.7%.

How many websites have an llms.txt file?

15.4% of the measured homepages publish llms.txt. SaaS sites lead at 28.2%, local businesses follow at 22.6% and government and education sites trail at 5.5%.

How much of a typical homepage is readable text?

The median homepage is 3.6% visible text by bytes. 61.5% are under 5% and only 15.6% reach 10%.

How many websites have no sitemap?

28.2% of measured sites have no sitemap a crawler can find, either declared in robots.txt or at the standard locations. Government and education sites are worst at 33.5%.

How was the data de-duplicated?

Each domain counts once, at its latest reading. 208,295 distinct domains were read; the 139,788 that could be measured end to end are the base for every percentage.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works