CrawlCheck

Findings · 2026-08-14 · By · 0 views

GPTBot, ClaudeBot, PerplexityBot: what each AI crawler is actually for

The eight most-blocked AI crawlers on the web sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent choices.

Blocking “AI crawlers” is not one decision. Eight named crawlers belong to different companies doing different things: some train models, some fetch pages to answer a question a user just asked, some index for search. Blocking the training crawler and the answer crawler from the same company are separate choices with separate consequences, and robots.txt is where you make them individually.

Share of all scans carrying each finding named aboveROBOTS_RULES_SHADOWED6.8%Share of all scans carrying eachfinding named aboveROBOTS_RULES_SHADOWED6.8%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Here is a number worth sitting with. Across the technology-detection population published by BuiltWith, read on 12 August 2026, the eight most commonly blocked AI crawlers are separated by about 12% end to end — roughly 96,000 sites at the bottom of that group and roughly 108,000 at the top. Apple’s crawler, OpenAI’s, Common Crawl’s, Anthropic’s, Amazon’s, ByteDance’s and Google’s training agent are all inside that band.

Below the band it falls off a cliff. Meta’s crawler sits near 69,000; ordinary Googlebot blocking sits near 42,000.

These are counts from a third party, quoted as reference points and attributed with the date we read them. They are not our dataset and we do not republish their table.

What near-lockstep actually means #

Eight companies with different products, different reputations and different relationships to publishers do not get blocked within 12% of each other by eight independent decisions. That is one copy-pasted block, one plugin default, one hosting toggle, applied by a lot of people who never compared the eight.

Which means most of what looks like AI policy on the web is inherited, not authored. Someone chose it once. Everyone else received it.

Eight operators, one spread #

CrawlerSites blocking it (BuiltWith, read 12 August 2026)
The eight most-blocked AI crawlerswithin about 12% of each other, end to end
Meta’s crawlernear 69,000
Ordinary Googlebotnear 42,000

Eight companies with different products and different relationships to publishers do not land within 12% of each other by eight independent judgements.

Three permissions, routinely collapsed into one #

Whether to allow AI training on your content is a real decision with real arguments on both sides, and it is not ours to make. But it is only a decision if it was made. An inherited block is not a position — it is a default nobody read.

And the three permissions get collapsed. Training, indexing and a live fetch during a conversation are separate things, often separate user agents, and people routinely block all three by accident while intending one. Blocking a training crawler does not remove you from answers. Blocking an indexing crawler does. Blocking a live-fetch agent means an assistant cannot open your page when a user asks it to — the one case where somebody specifically wanted you.

The separation is not theoretical; the operators publish it. OpenAI sends three distinct agents — one for the training and index crawl, one for the search index behind answers, one that fetches a page because a user asked for it in a chat. Anthropic sends three on the same split. Perplexity sends two. A single blanket rule hits all of them, and the one people least often mean to block is the third kind, where a person has already decided they want your page.

Two names in that list are not crawlers at all. Google-Extended and Applebot-Extended are permission tokens — they fetch nothing, ever. They exist so a site can decline training while remaining in search, which is the exact distinction most blanket rules destroy. Disallowing Applebot removes you from Siri and Apple search results; disallowing Applebot-Extended removes you from training and leaves search alone. Those two lines are one character apart in a file and opposite in effect. Every named agent, its operator and its purpose is listed separately for this reason.

The rule that silently deletes your other rules #

There is a mechanical trap underneath all of this, and it is the most common finding of its kind in our own panel. A crawler reads only the User-agent group that matches it best, and ignores every other group — including *. So naming an agent and giving it Allow: / does not add a permission on top of your general rules. It replaces them, for that agent, entirely.

Which means the well-meant edit — someone reads an article about AI visibility, adds User-agent: GPTBot and Allow: / to be welcoming — deletes every Disallow under * for that crawler. The staging paths, the internal search results, the checkout pages, the admin directories: all open, to that agent only. The file still parses. Every validator is happy. The agent still reaches the homepage, so nothing looks wrong on inspection.

We report it as ROBOTS_RULES_SHADOWED, and it currently sits on 28% of the sites this instrument measures daily — with a gap of zero between “ever” and “now”, because a file either has the defect or it does not and no Tuesday changes that. It is the most common robots-level finding in the panel, and it arrives from people trying to be more open rather than less.

And then the edge disagrees with the file #

A robots.txt is a request. What your edge actually does when a named crawler arrives is a separate fact, and the two disagree more often than anyone expects. In the panel, sites are found refusing an answer engine at the edge while their robots.txt says nothing of the kind — a firewall rule, a bot-management default, a rate limiter that only fires under load. Nothing in the file mentions it, so no amount of reading the file will find it.

That is why a scan here fetches the homepage as fifteen named identities from one address within one second, and compares every reply. Same second, same address: a difference between two of those replies is your edge deciding, not the internet being busy. And a claimed identity in your logs is not an identity either — nearly three in ten checkable claims here come from outside the ranges the named operator publishes, so a log full of one crawler’s name is not evidence about that crawler.

The check #

Two minutes: read your own robots.txt and ask, for each blocked agent, whether you would defend that line if a customer asked about it. If the answer is I did not know it was there, it was inherited. That is fine — but now it is a choice.

Then check the shape of the file, not just its contents. Every named group you have added is a complete replacement of * for that agent, so each one needs its own full set of Disallow lines repeated inside it. And if you meant to decline training only, the two permission tokens are the lines to use — not the crawler names beside them. The line-by-line version of that is in the guides.

A scan here lists every named crawler separately, with what your edge actually served each one, which is frequently not what your robots.txt says.

The three permissions, named #

Training crawlers build models; blocking one is a policy choice and this scanner does not score it. Search-index crawlers decide whether you appear in results at all. User-triggered fetchers open your page live when a person asks an assistant about you. Most robots.txt files that block AI block all three with one user-agent group, and the last one is the case where someone specifically wanted you. How to block AI crawlers lists which agent is which; do AI crawlers respect robots.txt covers what happens after you decide.

The defect that hides inside an allowlist #

One more thing to read for in your own file: if it names specific agents with their own groups, every rule in the * group stops applying to them. A crawler obeys only its most specific group. Sites that add an allowlist for search and AI agents routinely void every Disallow they thought still protected /wp-admin/ or their feeds. The scanner reports this as rules shadowed by a more specific group, and it was found on our own sites first.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Why do eight AI crawlers get blocked at similar rates?

Because the blocking decision is usually not made per crawler. Eight independent operators sitting within 12% of each other points at one shared list or default being applied, not eight separate judgements.

Does blocking AI crawlers require a per-operator decision?

In practice most sites inherit a default from a plugin, host or firewall vendor. The result is a policy nobody at the site chose and few can explain.

What should I check on my own site?

Whether your robots.txt names each agent deliberately, and whether the block you think you set is actually the one your edge is applying.

Can I allow answer engines but block training?

Yes, and that is the most common intent. Several operators run separate user-agents for training and for answering, so a group allowing one and disallowing the other is coherent policy rather than a mistake. The trap is that naming an agent in its own group voids the wildcard rules for it entirely.

Does blocking a crawler remove me from that product?

Sometimes, and not always in the way expected. Blocking a training crawler does not remove pages already in a model, and blocking an answer crawler can remove you from citations while leaving you in the underlying index. The two effects have different timelines.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works