Findings · 2026-10-06 · By VSNARY | Emmanuel Orta · 0 views
Only 6% of websites block an AI crawler: the full breakdown
Which AI crawlers 139,788 sites disallow in robots.txt, how training and search agents are treated differently, and how blocking varies by kind of site.
Across 139,788 homepages measured on 3-6 October 2026, 6.3% disallow at least one named AI crawler in robots.txt. GPTBot is blocked most (3.7%), then CCBot (3.1%), Bytespider (3.0%) and ClaudeBot (2.7%). Search and user-triggered agents are blocked far less: OAI-SearchBot 0.9%, Claude-User 0.5%. Publisher-heavy sites block most; government and local-business sites block least.
How many websites actually block AI crawlers? In the CrawlCheck dataset for 3-6 October 2026, the answer is 6.3%: 8,757 of 139,788 measured homepages disallow at least one named AI crawler in robots.txt. Every other site lets them all in, by choice or by default. The figures below count each domain once, at its latest reading, and name no site.
Crawler by crawler #
| Crawler (user-agent) | What it is for | Sites that block it |
|---|---|---|
| GPTBot | Collects pages for model training | 3.7% |
| CCBot | Builds an open web crawl widely used for training | 3.1% |
| Bytespider | Collects pages for model training | 3.0% |
| ClaudeBot | Collects pages for model training | 2.7% |
| Google-Extended | A control token for AI use of pages Google already crawls | 2.0% |
| ChatGPT-User | Fetches a page when a user asks for it | 1.8% |
| anthropic-ai | An older token for the same purpose as ClaudeBot | 1.8% |
| Applebot-Extended | A control token for AI use of pages Apple already crawls | 1.8% |
| meta-externalagent | Collects pages for AI use | 1.7% |
| PerplexityBot | Indexes pages for an answer engine | 1.6% |
| OAI-SearchBot | Indexes pages for search answers | 0.9% |
| Claude-User | Fetches a page when a user asks for it | 0.5% |
The AI crawlers page lists each of these with what it does and whether its operator publishes the addresses it crawls from.
Training is blocked more than search #
The pattern in the table is consistent: crawlers that collect pages for training are blocked two to five times as often as the agents that fetch a page to answer a question or build a search index. 3.1% of sites block GPTBot while leaving OAI-SearchBot open, which is the coherent version of an AI policy: stay out of the training data, stay in the answers. Blocking the search or user agents is what removes a site from the answers, and very few sites do it.
The control tokens are a separate case. Google-Extended and Applebot-Extended do not crawl anything themselves; they tell an existing search crawler whether its pages may be used for AI. Blocking them leaves search crawling untouched, which is why they are blocked at about the same rate as the training crawlers rather than as rarely as the search agents.
Who blocks #
| Lane | Homepages measured | Block any AI crawler | Block GPTBot |
|---|---|---|---|
| Everything else (includes publishers and media) | 89,780 | 8.1% | 5.0% |
| SaaS and software | 19,982 | 4.0% | 1.8% |
| US local businesses | 20,042 | 2.3% | 0.8% |
| Government and education | 9,984 | 2.1% | 1.4% |
Blocking concentrates where content is the product. The everything-else lane, which holds most publishers, blocks at four times the rate of government and local-business sites. A local business has every reason to be quoted and almost no content an AI company would pay for, and its robots.txt reflects that, mostly because nobody changed the default.
Bot blocks that are not AI blocks #
Many robots.txt files block bots that have nothing to do with AI: SEO crawlers, archive crawlers and regional search engines. 15.4% of measured sites disallow at least one named crawler of any kind, more than twice the AI rate. Some of that is a platform default rather than a choice: on one large website builder, 94% of sites block the same regional search crawler and almost none block an AI crawler. Counting every bot block as an AI block would more than double the figure, which is why the 6.3% above counts only AI crawlers.
What blocking does and does not do #
A disallow in robots.txt is a request. The major AI crawlers say they honour it, and the ones that publish the addresses they crawl from can be checked; the robots.txt testing guide shows how to confirm a rule reads the way you intended. It does not remove anything already collected, and it does not stop a crawler that ignores robots.txt or claims another name. If you block a crawler for some paths rather than all of them, its group has to repeat every wildcard rule you still want applied, as our correction on that check explains.
For the rest of the October figures, from llms.txt adoption to how much of a homepage is text, see the full dataset report.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What percentage of websites block GPTBot?
3.7% of 139,788 homepages measured on 3-6 October 2026 disallow GPTBot in robots.txt. It is the most blocked AI crawler in the dataset.
Do websites block AI search crawlers too?
Rarely. OAI-SearchBot is blocked by 0.9% of measured sites and Claude-User by 0.5%, against 2.7-3.7% for the main training crawlers.
Which kinds of sites block AI crawlers most?
Publisher-heavy sites. The lane that includes publishers and media blocks at 8.1%, against 2.1-2.3% for government, education and local-business sites.
Does blocking GPTBot remove my site from ChatGPT answers?
Not on its own. GPTBot collects training data; answers that cite pages rely on search and user-triggered agents such as OAI-SearchBot and ChatGPT-User, which have their own robots.txt groups.
Is blocking Google-Extended the same as blocking Googlebot?
No. Google-Extended is a control token that tells Google whether pages it already crawls may be used for AI. Blocking it leaves search crawling untouched.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.