CrawlCheck

Findings · 2026-08-14

Eight AI crawlers, eight separate companies, and one decision

The eight most-blocked AI crawlers on the web sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent choices.

Here is a number worth sitting with. Across the technology-detection population published by BuiltWith, read on 12 August 2026, the eight most commonly blocked AI crawlers are separated by about 12% end to end — roughly 96,000 sites at the bottom of that group and roughly 108,000 at the top. Apple’s crawler, OpenAI’s, Common Crawl’s, Anthropic’s, Amazon’s, ByteDance’s and Google’s training agent are all inside that band.

Below the band it falls off a cliff. Meta’s crawler sits near 69,000; ordinary Googlebot blocking sits near 42,000.

These are counts from a third party, quoted as reference points and attributed with the date we read them. They are not our dataset and we do not republish their table.

What near-lockstep actually means

Eight companies with different products, different reputations and different relationships to publishers do not get blocked within 12% of each other by eight independent decisions. That is one copy-pasted block, one plugin default, one hosting toggle, applied by a lot of people who never compared the eight.

Which means most of what looks like AI policy on the web is inherited, not authored. Someone chose it once. Everyone else received it.

Why that matters more than the ethics argument

Whether to allow AI training on your content is a real decision with real arguments on both sides, and it is not ours to make. But it is only a decision if it was made. An inherited block is not a position — it is a default nobody read.

And the three permissions get collapsed. Training, indexing and a live fetch during a conversation are separate things, often separate user agents, and people routinely block all three by accident while intending one. Blocking a training crawler does not remove you from answers. Blocking an indexing crawler does. Blocking a live-fetch agent means an assistant cannot open your page when a user asks it to — the one case where somebody specifically wanted you.

The check

Two minutes: read your own robots.txt and ask, for each blocked agent, whether you would defend that line if a customer asked about it. If the answer is I did not know it was there, it was inherited. That is fine — but now it is a choice.

A scan here lists every named crawler separately, with what your edge actually served each one, which is frequently not what your robots.txt says.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

All findings · The dataset · How the dataset works