Open dataset
What crawlers receive, measured across 1,629 scans
In one sentence
The CrawlCheck dataset is the aggregate of every scan run through the checker: how often each finding occurs (stale cache, no llms.txt, pages that are mostly code, AI opt-outs, broken sitemaps), the share of delivered bytes that is readable text, machine-file adoption, and how many named crawlers are refused — recomputed on every request, never naming a scanned domain, citable with the BibTeX and RIS records on the page.
Every scan run through the checker is counted here. No domains are published — only how often each defect occurs. Updated continuously.
1,629
scans counted
45
distinct finding codes seen
21%
carry the most common finding (STALE_CACHE_SERVED)
21.1%
carry a high or critical finding
Bar length is how common a finding is; bar colour is how serious — green is low severity, red is critical. The two are independent: the most common defect here is medium, and the critical ones are rare.
| Finding | Sites | Share | |
|---|---|---|---|
STALE_CACHE_SERVED | 347 | 21.3% | |
NO_LLMS_TXT | 325 | 20.0% | |
PAGE_IS_MOSTLY_CODE | 244 | 15.0% | |
AI_OPTOUT_SET | 156 | 9.6% | |
NAP_UNDECLARED_LISTING | 124 | 7.6% | |
NO_SITEMAP_FOUND | 89 | 5.5% | |
ROBOTS_NO_SITEMAP | 86 | 5.3% | |
ROBOTS_RULES_SHADOWED | 65 | 4.0% | |
ENTITY_COLLISION | 58 | 3.6% | |
ANSWER_ENGINE_REFUSED | 53 | 3.3% | |
ROBOTS_NOT_200 | 48 | 2.9% | |
HSTS_MISSING | 45 | 2.8% | |
UNIFORM_REFUSAL | 39 | 2.4% | |
SITEMAP_NO_LASTMOD | 31 | 1.9% | |
DECLARED_SITEMAP_BROKEN | 30 | 1.8% | |
SITEMAP_BLOCKED | 28 | 1.7% | |
HOST_DUPLICATE_200 | 26 | 1.6% | |
ENTITY_NO_COORDINATES | 26 | 1.6% | |
ROBOTS_DISALLOW_ALL | 26 | 1.6% | |
SITEMAP_IS_HTML | 20 | 1.2% | |
CHALLENGE_SERVED_200 | 20 | 1.2% | |
BROKEN_INTERNAL_LINK | 18 | 1.1% | |
MACHINE_FILE_CACHE_STALE | 17 | 1.0% | |
CHALLENGE_PINNED_AT_EDGE | 15 | 0.9% | |
ROBOTS_BLOCKED | 10 | 0.6% | |
HOMEPAGE_REDIRECTS_OFF_HOST | 9 | 0.6% | |
RENDER_RESOURCE_BLOCKED | 9 | 0.6% | |
HSTS_PRELOAD_INELIGIBLE | 7 | 0.4% | |
ORIGIN_BLOCKED_BEHIND_CACHE | 7 | 0.4% | |
NAP_NAME_DRIFT | 7 | 0.4% | |
CRAWLER_SERVED_LESS | 4 | 0.2% | |
ROBOTS_IS_HTML | 4 | 0.2% | |
CSP_REPORTING_BROKEN | 3 | 0.2% | |
CONTENT_NEEDS_JAVASCRIPT | 3 | 0.2% | |
STAGING_LEAK | 2 | 0.1% | |
SITEMAP_UNREADABLE | 2 | 0.1% | |
HOMEPAGE_REFUSED | 2 | 0.1% | |
ROBOTS_CTYPE | 2 | 0.1% | |
MACHINE_FILE_CACHE_SPLIT | 2 | 0.1% | |
ORIGIN_REFUSED_SCANNER | 2 | 0.1% | |
HSTS_PRELOAD_REMOVED | 2 | 0.1% | |
SITEMAP_ORIGIN_ERROR | 1 | 0.1% | |
DOMAIN_IS_ALIAS | 1 | 0.1% | |
HOMEPAGE_IS_INTERSTITIAL | 1 | 0.1% | |
DECLARED_SITEMAP_REFUSED | 1 | 0.1% |
How that compares
A share is only meaningful against a denominator. These are population counts for the same machine-layer signals, published by BuiltWith and read on 12 August 2026. They are counts of sites where the signal was detected anywhere on the web, not a sample of ours, and they are quoted here as reference points rather than reproduced as a dataset.
| Signal | Sites, web-wide | What it means |
|---|---|---|
| AppleBot disallowed | 108,452 | The largest single AI opt-out group |
| GPTBot disallowed | 107,182 | Within 1.2% of AppleBot |
| Common Crawl disallowed | 102,665 | The corpus most training sets draw on |
| ClaudeBot disallowed | 101,917 | Within 6% of AppleBot |
llms.txt published | 70,213 | The content map, still a minority signal |
| UCP profile published | 41,525 | Agentic-commerce discovery, Jan 2026 onward |
The eight largest opt-out groups sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent decisions, so most of what looks like AI policy on the web is one copy-pasted block, one plugin default or one hosting toggle — inherited rather than authored. That is the claim this dataset exists to test, one scan at a time.
Which structured facts actually exist
The same source counts detections of individual schema types. Declaring an entity is common; declaring the specific facts an assistant is asked for is an order of magnitude rarer.
| Schema type | Sites, web-wide | The question it answers |
|---|---|---|
| Organization | 354,358 | who this is |
| PostalAddress | 108,730 | where to send someone |
| ContactPoint | 91,622 | how to reach them |
| Offer | 55,182 | what it costs |
| Product | 46,455 | what is being sold |
| GeoCoordinates | 27,993 | where it is on a map |
| LocalBusiness | 26,108 | that it is a place of business |
| OpeningHoursSpecification | 22,972 | whether they are open now |
Fewer than one site in fifteen that declares an Organization declares opening hours. “Are they open?” and “how much is it?” are among the most common things anyone asks an assistant about a business, and the fields that answer them are the ones almost nobody publishes.
These are independent detection counts, not a joint distribution. Each row is the number of sites where that type was found anywhere; nothing here says which sites overlap, and no coverage rate should be inferred by dividing one row by another. Counts read from BuiltWith on 13 August 2026.
What a single crawl cannot see
77.1% of measured sites carry a defect today. 95.8% carried one at least once between 2026-08-14 and 2026-09-11. 9 look clean now and did not for at least one day in that window.
| Across 48 sites measured every day | Share |
|---|---|
| Carry a defect today | 77.1% |
| Carried one at least once in the window | 95.8% |
| The gap | 18.7 points |
| Look clean now and did not, at least once | 9 |
Both numbers describe the same sites over 29 days, 2026-08-14 to 2026-09-11. A one-off crawl at any budget can only ever produce the first row — the second needs the same sites to have been measured before the question was asked. Sites added after the window opened are excluded (208) because they have had fewer chances to be seen broken. No scanned domain is named here or anywhere else on this site.
Arrivals across the network
measured on 3 sites — too few to describe as a population, so no rate is published
Forged crawler identities
23% of the requests that named themselves as a known crawler here were not that crawler. 1369 of 5940 checkable claims came from an address outside the range the operator publishes.
Forged rate by claimed identity — share of checkable claims from outside the operator’s published range
| Claimed to be | Claims | Verified | Forged | Unverifiable | Forged rate |
|---|---|---|---|---|---|
| Meta-ExternalAgent | 1479 | 0 | 0 | 1479 | no feed |
| Googlebot | 1199 | 1128 | 71 | — | 5.9% |
| PerplexityBot | 1041 | 903 | 138 | — | 13.3% |
| GPTBot | 910 | 726 | 184 | — | 20.2% |
| Applebot | 851 | 629 | 222 | — | 26.1% |
| ClaudeBot | 748 | 646 | 102 | — | 13.6% |
| AhrefsBot | 574 | 0 | 0 | 574 | no feed |
| Amazonbot | 514 | 0 | 0 | 514 | no feed |
| ChatGPT-User | 375 | 197 | 178 | — | 47.5% |
| OAI-SearchBot | 345 | 216 | 129 | — | 37.4% |
| Baiduspider | 308 | 0 | 0 | 308 | no feed |
| YandexBot | 297 | 0 | 0 | 297 | no feed |
A user-agent is a claim, not an identity. Verified means the source IP sits inside a range the operator publishes — OpenAI, Anthropic, Google, Microsoft, Perplexity and Apple all publish one. Unverifiable is not forgery: some operators publish no range at all, so their requests can be neither confirmed nor accused, and they are excluded from the rate rather than counted against it. This is traffic to this site only, and this site is small — it is a floor on the problem, not a survey of the web.
What is in the sample
Coverage matters as much as the counts: a defect rate measured entirely on one platform is a fact about that platform. The scanner deliberately widens the sample rather than scanning the same kind of site repeatedly, and the queue below is what it has not reached yet.
Crawlers we have actually seen
Everything else on this page comes from probes — we send a request wearing a crawler’s name and record what comes back. This table is the opposite: visits nobody asked for, on 4 sites we operate, over 30 days. Verified means Cloudflare matched the request to the address ranges that crawler’s operator publishes.
| Crawler | Visits | Verified | Unverified | Most-requested path |
|---|---|---|---|---|
| AhrefsBot | 3055 | 3037 | 18 | /entitymap-sitemap.xml |
| ClaudeBot | 2818 | 2445 | 373 | /entitymap-sitemap.xml |
| Applebot | 2366 | 1991 | 375 | / |
| Googlebot | 2353 | 1946 | 407 | / |
| Amazonbot | 2013 | 1528 | 485 | /wp-json/oembed/1.0/embed |
| GPTBot | 1933 | 1244 | 689 | / |
| meta-externalagent | 1879 | 1512 | 367 | / |
| PerplexityBot | 1357 | 0 | 1357 | / |
| Bingbot | 1189 | 964 | 225 | / |
| OAI-SearchBot | 877 | 501 | 376 | /robots.txt |
| SemrushBot | 856 | 843 | 13 | /robots.txt |
| ChatGPT-User | 728 | 210 | 518 | / |
| Bytespider | 374 | 58 | 316 | /robots.txt |
| Claude-User | 204 | 47 | 157 | /robots.txt |
| Google-Extended | 199 | 0 | 199 | /fetch |
| CCBot | 193 | 0 | 193 | / |
| Perplexity-User | 148 | 0 | 148 | /config/secrets.yml |
| YouBot | 34 | 21 | 13 | /robots.txt |
| DuckAssistBot | 16 | 11 | 5 | /robots.txt |
| cohere-ai | 4 | 0 | 4 | / |
| Applebot-Extended | 2 | 0 | 2 | / |
| anthropic-ai | 2 | 0 | 2 | / |
| Diffbot | 1 | 0 | 1 | / |
How to read this, and how not to. 22601 visits across a handful of sites is a sample, not a census — it says what reached these sites, not what the web receives. And unverified is not proof of an impostor: plenty of legitimate traffic arrives from ranges nobody publishes, and our own testing appears in these counts as unverified because it is. What the column does show is that the name in a user-agent string is a claim, and it is checkable.
| Dimension | Observed |
|---|---|
| Platform | wordpress 235 · unknown 935 · nextjs 148 · webflow 27 · drupal 43 · shopify 17 · wix 7 · squarespace 11 |
| Rendering | server 1115 · client 308 |
| Size | medium 351 · small 146 · large 8 |
| Business type | local-service 6 · other 1006 · publisher 284 · ecommerce 41 · gov-edu 19 · saas 51 · docs 16 |
| Language | en 891 · non-en 373 |
| Queued, not yet scanned | 1,981 domains |
Platform, rendering, size and business type are read from what each site discloses in its own response. Most sites now answer from behind a CDN, which replaces the origin identity in the headers, so unknown here means the site did not disclose one — not that we did not look. We would rather publish a large unknown than a confident guess.
How to read these numbers
Each row counts sites, not pages, and a site is counted once per finding however many
times it was scanned. A finding is recorded only when the check actually returned an answer:
lookups that failed on our side are excluded rather than counted as clean, which is why the totals
here can be smaller than the scan count. Domains are never named, in aggregate or individually,
and any operator can exclude a domain from this dataset in one line of robots.txt
— see the policy page.
Every figure on this page is recomputed when the page is requested, so a number quoted
elsewhere is a snapshot of the moment it was read. The same counts are served as JSON at
/api/public/counts with an at
timestamp — no key, CORS open — so anyone quoting us can check the current value
instead of taking ours. A record is only worth more than a single reading if the reading
can be repeated.
Cite this dataset
Every figure here is recomputed from the record when you load the page, so a citation should carry the access date as well as the year. Both records below do.
BibTeX download
@online{crawlcheck_dataset_2026,
author = {{CrawlCheck}},
organization = {CrawlCheck},
type = {Live dataset},
language = {english},
keywords = {AI crawlers, answer engines, llms.txt, robots.txt, machine readability},
abstract = {Aggregate measurements of what answer engines and AI crawlers receive from public websites - machine-file adoption, crawler access outcomes and payload composition, recomputed continuously across the scanned corpus.},
title = {Dataset --- {CrawlCheck}},
year = {2026},
url = {https://crawlcheck.io/data},
urldate = {2026-09-11},
note = {Continuously updated; figures recomputed on each request}
}
RIS download
TY - DATA TI - Dataset — CrawlCheck AU - CrawlCheck PB - CrawlCheck KW - AI crawlers KW - Answer engines KW - Machine readability AB - Aggregate measurements of what answer engines and AI crawlers receive from public websites - machine-file adoption, crawler access outcomes and payload composition, recomputed continuously across the scanned corpus. PY - 2026 UR - https://crawlcheck.io/data Y2 - 2026/09/11 N1 - Continuously updated; figures recomputed on each request ER -
The RIS type is DATA rather than ELEC — this is a dataset, and reference managers file the two differently.