Findings · 2026-10-07 · By VSNARY | Emmanuel Orta · 0 views
Eight weeks of AI agents on one site: who was really who, and 47,700 visits nobody can verify
Every request that claimed to be a named crawler on crawlcheck.io since 13 August, checked against the address ranges its operator publishes. The big crawlers are who they say 97 to 99 times in 100. The browsing agents are forged far more often. Three names arrive only as forgeries. And the busiest visitors of all cannot be verified by anyone.
On crawlcheck.io between 13 August and 7 October 2026, 24,627 requests claiming a named AI or search crawler came from the address ranges that crawler's operator publishes, and 787 did not. Forgery is rare for the index crawlers (Googlebot 1.3%, PerplexityBot 1.4%, ClaudeBot 1.9%, GPTBot 2.4%, Applebot 2.5%, Bingbot 3.7%, OAI-SearchBot 5.0%) and common for the user-triggered fetchers (ChatGPT-User 12.4%, Claude-User 40.2%, Perplexity-User 86.2%). Google-Extended (49 claims), Claude-SearchBot (37) and anthropic-ai (4) were never verified once: they are robots.txt tokens, not crawlers, so every request wearing them is a forgery. Every verified agent respected robots.txt, and none exceeded 30 requests a minute. 47,718 further visits used 17 names whose operators publish no address ranges, so nothing they did can be told from an impersonator.
A request arrives at your site with GPTBot in its user agent. Is it OpenAI? The string is a claim anyone can type. The only check that settles it is the source address against the ranges the operator publishes, and most sites never make it. This one does, for every request, and has since 13 August 2026. The record is public and signed at /api/v1/conduct; this post reads it on 7 October.
One property, eight weeks, 25,414 requests that named a verifiable agent. It is a census of this site, not a sample of the web, and that is the point: these are the agents that actually came, counted one by one.
The index crawlers are who they say #
| Agent | Operator | Verified requests | Forged claims | Forged share |
|---|---|---|---|---|
| Googlebot | 5,514 | 74 | 1.3% | |
| PerplexityBot | Perplexity | 4,979 | 71 | 1.4% |
| GPTBot | OpenAI | 4,013 | 99 | 2.4% |
| ClaudeBot | Anthropic | 3,517 | 67 | 1.9% |
| Applebot | Apple | 3,212 | 83 | 2.5% |
| OAI-SearchBot | OpenAI | 1,486 | 79 | 5.0% |
| Bingbot | Microsoft | 1,226 | 47 | 3.7% |
Between one and five claims in a hundred are forgeries. The other 95 to 99 are the real crawler, from the real ranges. For a site deciding whether to serve, block or meter a crawler by name, that is the number that matters: the name is a reliable signal for these seven, provided the address is checked. Unchecked, 1 to 5% of what you attribute to Google or OpenAI is someone else wearing their coat.
The browsing agents are forged far more #
| Agent | Operator | Verified | Forged | Forged share |
|---|---|---|---|---|
| ChatGPT-User | OpenAI | 508 | 72 | 12.4% |
| Claude-User | Anthropic | 79 | 53 | 40.2% |
| Perplexity-User | Perplexity | 8 | 50 | 86.2% |
These are the agents that fetch a page because a person asked, in a chat, right now. They are also the names a scraper picks when it wants to look like a person’s assistant rather than a crawler, because sites that block GPTBot often let ChatGPT-User through. On this site, more than one Claude-User request in three and more than four Perplexity-User requests in five did not come from the operator. If your access policy treats user-triggered agents more generously than crawlers, that policy is being used, and the address check is the only thing standing between it and anyone who read the docs.
Three names that only ever arrive as forgeries #
Google-Extended was claimed 49 times and verified never. Claude-SearchBot, 37 times, never. anthropic-ai, 4 times, never. Two of the three are not crawlers that happen to be spoofed. Google-Extended and anthropic-ai are robots.txt tokens: names a site writes in a rule to control training use, which no request ever carries, because the operator’s crawler identifies as Googlebot or ClaudeBot and reads the token in the file. A request that puts the token in its user agent was written by someone who learned the names from a robots.txt guide and not from the operator. It is a forgery by construction, and a cheap one to drop at the edge. Claude-SearchBot is different: a real agent name Anthropic publishes ranges for. In eight weeks the real one never came, and 37 requests wearing its name came from somewhere else.
What the verified agents did #
Identity is half the record. The other half is conduct, and here it is clean. Across every verified agent with enough requests to judge — 24,000 of them — no verified request fetched a path this site’s robots.txt disallows, including the paths disallowed in a group that names the agent specifically. No verified agent exceeded this site’s own limit of 30 requests a minute; the busiest minute belonged to Googlebot at 3. The crawlers people worry about, when they are really themselves, read the rules and kept to them on this site for eight weeks.
What that does not say: whether they respected anyone else’s rules, or what they did with the pages. One property, one robots.txt. But it is a measured baseline where there was none, and any site running the same check can publish its own.
47,700 visits that nobody can verify #
The largest visitors to this site are not in either table, because they cannot be. Seventeen names arrived without a published address range to check them against:
| Name claimed | Requests | Why it cannot be verified |
|---|---|---|
| Amazonbot | 20,610 | The operator publishes no machine-readable address ranges. Every request is unverifiable: the real agent and an impersonator look identical. |
| Meta-ExternalAgent | 9,665 | |
| SemrushBot | 6,835 | |
| AhrefsBot | 4,590 | |
| YandexBot | 2,334 | |
| DataForSeoBot | 2,229 | |
| Baiduspider, Bytespider, YouBot, CCBot, DuckDuckBot and six more | 1,455 |
Amazonbot alone made more requests here than Googlebot and GPTBot together, and not one of them can be attributed to Amazon rather than to anyone who typed the name. That is not an accusation; the traffic may be entirely genuine. It is a gap: a site cannot allow a crawler it cannot recognise without allowing everyone who imitates it, so the honest options are to block the name or to serve everyone who uses it. Publishing a range file is a few lines of JSON. The seven operators in the first table did it, and it is why their agents can be let in by name.
Run the same check on your own logs #
Nothing here needed a login or a key. Visitor Resolve takes an address and a user agent and answers verified, spoofed, unverifiable or not a known agent, against the ranges each operator publishes, one request or a hundred at a time. Sites behind Cloudflare can run it at the edge on every request with the published spec. The agent records above are updated continuously, signed, and verify offline with crawlcheck-verify.mjs.
Method #
Every request to crawlcheck.io whose user agent names a known agent is checked at arrival against the address ranges that agent’s operator publishes. Verified means inside the ranges; forged means the name with an address outside every range the operator publishes for it. Agents whose operators publish no ranges are counted but never judged. Conduct is assessed only for agents with at least 50 verified requests; robots.txt conduct counts verified fetches of paths disallowed for every agent including a group naming that one, and rate is the busiest minute of verified requests since 7 October, so both are lower bounds. Figures are as of 7 October 2026 and change daily at the endpoint linked above.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
How often is GPTBot spoofed?
On crawlcheck.io over eight weeks, 99 of 4,112 requests claiming GPTBot (2.4%) came from outside OpenAI's published ranges. Googlebot was forged 1.3% of the time, ClaudeBot 1.9%, Bingbot 3.7%, OAI-SearchBot 5.0%.
Is ChatGPT-User or Claude-User traffic real?
Mostly for ChatGPT-User (87.6% verified), much less for Claude-User (59.8% verified) and rarely for Perplexity-User (13.8% verified) on this site. User-triggered agent names are the ones scrapers borrow, because sites often let them through where they block crawlers.
Why can't Amazonbot or SemrushBot be verified?
Their operators publish no machine-readable list of address ranges, so no request carrying the name can be tied to the operator. The real agent and an impersonator are indistinguishable; a site can only allow or block the name as a whole.
Do AI crawlers respect robots.txt?
On this site, every verified agent did: 24,000 verified requests over eight weeks, none to a disallowed path, none above 30 requests a minute. That is one property's measurement, not a claim about the web; any site can publish its own with the same check.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.