CrawlCheck

Open dataset

What crawlers receive, measured across 1,629 scans

In one sentence

The CrawlCheck dataset is the aggregate of every scan run through the checker: how often each finding occurs (stale cache, no llms.txt, pages that are mostly code, AI opt-outs, broken sitemaps), the share of delivered bytes that is readable text, machine-file adoption, and how many named crawlers are refused — recomputed on every request, never naming a scanned domain, citable with the BibTeX and RIS records on the page.

Every scan run through the checker is counted here. No domains are published — only how often each defect occurs. Updated continuously.

1,629

scans counted

45

distinct finding codes seen

21%

carry the most common finding (STALE_CACHE_SERVED)

21.1%

carry a high or critical finding

Bar length is how common a finding is; bar colour is how serious — green is low severity, red is critical. The two are independent: the most common defect here is medium, and the critical ones are rare.

FindingSitesShare 
STALE_CACHE_SERVED34721.3%
NO_LLMS_TXT32520.0%
PAGE_IS_MOSTLY_CODE24415.0%
AI_OPTOUT_SET1569.6%
NAP_UNDECLARED_LISTING1247.6%
NO_SITEMAP_FOUND895.5%
ROBOTS_NO_SITEMAP865.3%
ROBOTS_RULES_SHADOWED654.0%
ENTITY_COLLISION583.6%
ANSWER_ENGINE_REFUSED533.3%
ROBOTS_NOT_200482.9%
HSTS_MISSING452.8%
UNIFORM_REFUSAL392.4%
SITEMAP_NO_LASTMOD311.9%
DECLARED_SITEMAP_BROKEN301.8%
SITEMAP_BLOCKED281.7%
HOST_DUPLICATE_200261.6%
ENTITY_NO_COORDINATES261.6%
ROBOTS_DISALLOW_ALL261.6%
SITEMAP_IS_HTML201.2%
CHALLENGE_SERVED_200201.2%
BROKEN_INTERNAL_LINK181.1%
MACHINE_FILE_CACHE_STALE171.0%
CHALLENGE_PINNED_AT_EDGE150.9%
ROBOTS_BLOCKED100.6%
HOMEPAGE_REDIRECTS_OFF_HOST90.6%
RENDER_RESOURCE_BLOCKED90.6%
HSTS_PRELOAD_INELIGIBLE70.4%
ORIGIN_BLOCKED_BEHIND_CACHE70.4%
NAP_NAME_DRIFT70.4%
CRAWLER_SERVED_LESS40.2%
ROBOTS_IS_HTML40.2%
CSP_REPORTING_BROKEN30.2%
CONTENT_NEEDS_JAVASCRIPT30.2%
STAGING_LEAK20.1%
SITEMAP_UNREADABLE20.1%
HOMEPAGE_REFUSED20.1%
ROBOTS_CTYPE20.1%
MACHINE_FILE_CACHE_SPLIT20.1%
ORIGIN_REFUSED_SCANNER20.1%
HSTS_PRELOAD_REMOVED20.1%
SITEMAP_ORIGIN_ERROR10.1%
DOMAIN_IS_ALIAS10.1%
HOMEPAGE_IS_INTERSTITIAL10.1%
DECLARED_SITEMAP_REFUSED10.1%

How that compares

A share is only meaningful against a denominator. These are population counts for the same machine-layer signals, published by BuiltWith and read on 12 August 2026. They are counts of sites where the signal was detected anywhere on the web, not a sample of ours, and they are quoted here as reference points rather than reproduced as a dataset.

SignalSites, web-wideWhat it means
AppleBot disallowed108,452The largest single AI opt-out group
GPTBot disallowed107,182Within 1.2% of AppleBot
Common Crawl disallowed102,665The corpus most training sets draw on
ClaudeBot disallowed101,917Within 6% of AppleBot
llms.txt published70,213The content map, still a minority signal
UCP profile published41,525Agentic-commerce discovery, Jan 2026 onward

The eight largest opt-out groups sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent decisions, so most of what looks like AI policy on the web is one copy-pasted block, one plugin default or one hosting toggle — inherited rather than authored. That is the claim this dataset exists to test, one scan at a time.

Which structured facts actually exist

The same source counts detections of individual schema types. Declaring an entity is common; declaring the specific facts an assistant is asked for is an order of magnitude rarer.

Schema typeSites, web-wideThe question it answers
Organization354,358who this is
PostalAddress108,730where to send someone
ContactPoint91,622how to reach them
Offer55,182what it costs
Product46,455what is being sold
GeoCoordinates27,993where it is on a map
LocalBusiness26,108that it is a place of business
OpeningHoursSpecification22,972whether they are open now

Fewer than one site in fifteen that declares an Organization declares opening hours. “Are they open?” and “how much is it?” are among the most common things anyone asks an assistant about a business, and the fields that answer them are the ones almost nobody publishes.

These are independent detection counts, not a joint distribution. Each row is the number of sites where that type was found anywhere; nothing here says which sites overlap, and no coverage rate should be inferred by dividing one row by another. Counts read from BuiltWith on 13 August 2026.

What a single crawl cannot see

77.1% of measured sites carry a defect today. 95.8% carried one at least once between 2026-08-14 and 2026-09-11. 9 look clean now and did not for at least one day in that window.

Across 48 sites measured every dayShare
Carry a defect today77.1%
Carried one at least once in the window95.8%
The gap18.7 points
Look clean now and did not, at least once9

Both numbers describe the same sites over 29 days, 2026-08-14 to 2026-09-11. A one-off crawl at any budget can only ever produce the first row — the second needs the same sites to have been measured before the question was asked. Sites added after the window opened are excluded (208) because they have had fewer chances to be seen broken. No scanned domain is named here or anywhere else on this site.

Arrivals across the network

measured on 3 sites — too few to describe as a population, so no rate is published

Forged crawler identities

23% of the requests that named themselves as a known crawler here were not that crawler. 1369 of 5940 checkable claims came from an address outside the range the operator publishes.

Forged rate by claimed identity — share of checkable claims from outside the operator’s published range

Google-Extended100%Claude-SearchBot100%Perplexity-User87.1%Claude-User63.9%Bingbot59.9%ChatGPT-User47.5%OAI-SearchBot37.4%Applebot26.1%GPTBot20.2%ClaudeBot13.6%
Claimed to beClaimsVerifiedForgedUnverifiableForged rate
Meta-ExternalAgent1479001479no feed
Googlebot11991128715.9%
PerplexityBot104190313813.3%
GPTBot91072618420.2%
Applebot85162922226.1%
ClaudeBot74864610213.6%
AhrefsBot57400574no feed
Amazonbot51400514no feed
ChatGPT-User37519717847.5%
OAI-SearchBot34521612937.4%
Baiduspider30800308no feed
YandexBot29700297no feed

A user-agent is a claim, not an identity. Verified means the source IP sits inside a range the operator publishes — OpenAI, Anthropic, Google, Microsoft, Perplexity and Apple all publish one. Unverifiable is not forgery: some operators publish no range at all, so their requests can be neither confirmed nor accused, and they are excluded from the rate rather than counted against it. This is traffic to this site only, and this site is small — it is a floor on the problem, not a survey of the web.

What is in the sample

Coverage matters as much as the counts: a defect rate measured entirely on one platform is a fact about that platform. The scanner deliberately widens the sample rather than scanning the same kind of site repeatedly, and the queue below is what it has not reached yet.

Crawlers we have actually seen

Everything else on this page comes from probes — we send a request wearing a crawler’s name and record what comes back. This table is the opposite: visits nobody asked for, on 4 sites we operate, over 30 days. Verified means Cloudflare matched the request to the address ranges that crawler’s operator publishes.

CrawlerVisitsVerifiedUnverifiedMost-requested path
AhrefsBot3055303718/entitymap-sitemap.xml
ClaudeBot28182445373/entitymap-sitemap.xml
Applebot23661991375/
Googlebot23531946407/
Amazonbot20131528485/wp-json/oembed/1.0/embed
GPTBot19331244689/
meta-externalagent18791512367/
PerplexityBot135701357/
Bingbot1189964225/
OAI-SearchBot877501376/robots.txt
SemrushBot85684313/robots.txt
ChatGPT-User728210518/
Bytespider37458316/robots.txt
Claude-User20447157/robots.txt
Google-Extended1990199/fetch
CCBot1930193/
Perplexity-User1480148/config/secrets.yml
YouBot342113/robots.txt
DuckAssistBot16115/robots.txt
cohere-ai404/
Applebot-Extended202/
anthropic-ai202/
Diffbot101/

How to read this, and how not to. 22601 visits across a handful of sites is a sample, not a census — it says what reached these sites, not what the web receives. And unverified is not proof of an impostor: plenty of legitimate traffic arrives from ranges nobody publishes, and our own testing appears in these counts as unverified because it is. What the column does show is that the name in a user-agent string is a claim, and it is checkable.

DimensionObserved
Platformwordpress 235 · unknown 935 · nextjs 148 · webflow 27 · drupal 43 · shopify 17 · wix 7 · squarespace 11
Renderingserver 1115 · client 308
Sizemedium 351 · small 146 · large 8
Business typelocal-service 6 · other 1006 · publisher 284 · ecommerce 41 · gov-edu 19 · saas 51 · docs 16
Languageen 891 · non-en 373
Queued, not yet scanned1,981 domains

Platform, rendering, size and business type are read from what each site discloses in its own response. Most sites now answer from behind a CDN, which replaces the origin identity in the headers, so unknown here means the site did not disclose one — not that we did not look. We would rather publish a large unknown than a confident guess.

How to read these numbers

Each row counts sites, not pages, and a site is counted once per finding however many times it was scanned. A finding is recorded only when the check actually returned an answer: lookups that failed on our side are excluded rather than counted as clean, which is why the totals here can be smaller than the scan count. Domains are never named, in aggregate or individually, and any operator can exclude a domain from this dataset in one line of robots.txt — see the policy page.

Every figure on this page is recomputed when the page is requested, so a number quoted elsewhere is a snapshot of the moment it was read. The same counts are served as JSON at /api/public/counts with an at timestamp — no key, CORS open — so anyone quoting us can check the current value instead of taking ours. A record is only worth more than a single reading if the reading can be repeated.

Run a scan

Cite this dataset

Every figure here is recomputed from the record when you load the page, so a citation should carry the access date as well as the year. Both records below do.

BibTeX download

@online{crawlcheck_dataset_2026,
  author       = {{CrawlCheck}},
  organization = {CrawlCheck},
  type         = {Live dataset},
  language     = {english},
  keywords     = {AI crawlers, answer engines, llms.txt, robots.txt, machine readability},
  abstract     = {Aggregate measurements of what answer engines and AI crawlers receive from public websites - machine-file adoption, crawler access outcomes and payload composition, recomputed continuously across the scanned corpus.},
  title   = {Dataset --- {CrawlCheck}},
  year    = {2026},
  url     = {https://crawlcheck.io/data},
  urldate = {2026-09-11},
  note    = {Continuously updated; figures recomputed on each request}
}

RIS download

TY  - DATA
TI  - Dataset — CrawlCheck
AU  - CrawlCheck
PB  - CrawlCheck
KW  - AI crawlers
KW  - Answer engines
KW  - Machine readability
AB  - Aggregate measurements of what answer engines and AI crawlers receive from public websites - machine-file adoption, crawler access outcomes and payload composition, recomputed continuously across the scanned corpus.
PY  - 2026
UR  - https://crawlcheck.io/data
Y2  - 2026/09/11
N1  - Continuously updated; figures recomputed on each request
ER  - 

The RIS type is DATA rather than ELEC — this is a dataset, and reference managers file the two differently.

Questions about this page

QDoes the dataset name the sites it measured?
No. It publishes counts and shares only. A domain owner can remove a site from the aggregates entirely with one line in robots.txt.
QWhat is the most common finding across sites?
The table on this page is recomputed on every request, so read the current top row there rather than a number quoted elsewhere. Common findings are a stale cached copy served to crawlers, no llms.txt, and pages where almost none of the delivered bytes are readable text.
QCan I cite this dataset?
Yes. BibTeX and RIS records are on the page, and the figures carry the date they were computed, because they change as scans accrue.
QIs this a random sample of the web?
No. It is the set of sites people chose to scan plus a seeded corpus, which skews toward sites somebody already suspected. Percentiles here describe this corpus, not the web, and the page says so.

Keep reading