Findings · 2026-09-12 · By VSNARY | Emmanuel Orta · 0 views
1,582 requests arrived under a name Google says it never sends
Google-Extended is a robots.txt control token. Google's own documentation says it has no HTTP user-agent string. Over seven days, 1,582 requests reached our zones carrying it anyway.
Between 5 and 11 September 2026 we read Cloudflare's own bot-verification verdict on every request that reached sixteen zones on one account. Not a guess from a user-agent string, and not an outside-in probe: the verdict the network that served the request already recorded. Two counts per crawler name, over the same requests, in the same window - how many Cloudflare verified, and how many carried the same name and were not verified.
Eleven names came back with a rate worth publishing. Five came back having never been verified once. One of those five should not exist as a user-agent at all.
The name that is not a crawler
Google-Extended is the robots.txt token that governs whether content Google has already crawled may be used to train and ground Gemini. It is a policy control, not a fetcher. Google's crawler documentation is explicit about it: Google-Extended does not have a separate HTTP request user agent string, crawling is done with existing Google user agent strings, and the robots.txt token is used in a control capacity.
So there is nothing to verify, because nothing is supposed to arrive. Over the window, 1,582 requests reached these zones carrying Google-Extended in the user-agent. Whatever sent them, it was not Google using a documented identity, because Google documents that this identity is never sent.
That is a different statement from the one our page was making. A rate of 0% verified reads as an accusation about an operator's fleet. The true statement here is narrower and stronger: a name that its owner says never appears on the wire appeared 1,582 times.
What the other four mean, which is not the same thing
PerplexityBot arrived 4,893 times and was verified zero times. That is not evidence that 4,893 requests were forged. Cloudflare removed Perplexity from its Verified Bots programme in August 2025, naming PerplexityBot and Perplexity-User as the identifiers. Once a crawler is outside the programme there is no verification path at all, so every request under that name lands unverified by construction - a fact about the reference data, not about the requests.
CCBot and Claude-SearchBot are in the same position from where we stand: never verified once, across sixteen zones and seven days. We have no external evidence about either, and we are not going to invent any. What we can say is that we have never seen a verification succeed for those names here, which is exactly the sentence the page now prints instead of a percentage.
The rates that are real
For the eleven names where verification demonstrably works, the spread is wide and it does not line up with how people talk about these crawlers.
| Crawler name | Verified | Same name, unverified | Share verified |
|---|---|---|---|
| SemrushBot | 4,212 | 1,309 | 76.3% |
| AhrefsBot | 30,270 | 14,493 | 67.6% |
| bingbot | 3,857 | 1,880 | 67.2% |
| meta-externalagent | 4,352 | 3,735 | 53.8% |
| Googlebot | 5,284 | 4,974 | 51.5% |
| ClaudeBot | 3,829 | 4,718 | 44.8% |
| GPTBot | 1,538 | 2,948 | 34.3% |
| PerplexityBot | 0 | 4,893 | withheld |
| Google-Extended | 0 | 1,582 | withheld |
Even Googlebot, the most verifiable crawler on the internet, sits at 51.5%. Roughly half the requests arriving at these zones under Googlebot's name are not Googlebot by Cloudflare's reckoning. That is the number most site owners would be most surprised by, and it is the one we very nearly failed to publish at all.
Our own endpoint was publishing a false zero
Before today, the public aggregate read the verified side out of a field that keeps only the crawlers Cloudflare files under a category beginning AI, and keyed them with a looser matcher than the unverified side used. The denominator counted every crawler we publish a pattern for; the numerator counted a subset of them, under different keys.
Dividing one into the other produced this: 19 of 30 rows read 0% verified, and Googlebot was one of them. Cloudflare files Googlebot as a search engine crawler, not as AI, so the numerator never counted a single one of its 5,284 verified requests. The same went for bingbot, AhrefsBot, SemrushBot, MJ12bot, Baiduspider and YandexBot - every one of them verified thousands of times, every one of them published at zero.
| Name | Published before | Published now |
|---|---|---|
| Googlebot | 0% | 51.5% |
| bingbot | 0% | 67.2% |
| AhrefsBot | 0% | 67.6% |
| SemrushBot | 0% | 76.3% |
| MJ12bot | 0% | 85.8% |
| Baiduspider | 0% | 53.1% |
Nobody was reading the endpoint yet, which is the only reason this is a note rather than an apology. It is still the worst kind of defect we ship: a number that looks measured, on a page whose entire proposition is that the numbers are measured. We have done this to ourselves before, in the other direction, and counted our own traffic as forged.
A rate needs a verification path to exist
Fixing the key space was the easy half. The harder half is that 0% says two opposite things and cannot tell you which. Either the crawler is verifiable and not one request passed, or there is no verification path for that name and nothing could ever pass. Printed as a percentage, a reader takes the first reading every time.
The distinguishing evidence was already in the data. A name verified even once, on any zone in the window, has a working path, and its rate means something. A name never verified across sixteen zones and seven days does not, and the honest output is the counts plus the reason - which is what those five rows now carry, alongside the operator's own published statement where one exists.
The field that records it is three-valued, not two: true where we have seen a verification succeed, and null where we have not. Never false. We do not know that Cloudflare cannot verify CCBot; we know we have not seen it happen. That distinction is the same one every unscored row on a report rests on.
What this does not claim
These are sixteen zones on one account, measured because we operate them, not a sample of the web. Unverified is not forged: a legitimate crawler behind a proxy, or one whose ranges Cloudflare has not adopted, lands in the same bucket as an impostor. The unverified side is read from the three hundred most common user-agent strings per zone per day, so a rare impostor on a busy zone can fall outside the list, and the response records the zone-days where that limit was hit.
What survives all of that is narrow and worth saying anyway. Half the traffic arriving as Googlebot is not Googlebot. And 1,582 requests introduced themselves with a name that the company that owns it says it never sends. If you want to see the current figures rather than these, they are on the crawler reference, recomputed from the same records.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
Related findings
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.