Findings · 2026-08-18 · By VSNARY | Emmanuel Orta · 1 view
Our AI-crawler forgery rate was counting our own browser extension
Our extension forges a crawler user-agent on every probe — by design. To this site's own telemetry that is indistinguishable from a stranger doing the same thing, so for weeks our users inflated the number we publish about everyone else.
A crawler-verification check that counts every failed verification as a forgery will report your own tooling as an attacker. Ours did: requests to our API and report URLs carrying a crawler name were adjudicated as spoofed, when they were people checking a report we had produced. Path matters — a failed verification on a content page is a different event from one on an API endpoint, and only the first is evidence of impersonation.
This site measures how often a request claiming to be a named AI crawler actually comes from that operator's published address range. The headline was a big number, and big numbers travel. This one was partly measuring us.
How the error worked #
Our browser extension exists to answer a question a server cannot: what does a crawler receive from your address, after your JavaScript runs. To do that it sends requests wearing a crawler's user-agent. That is the entire point of the tool, and it is honest work — but a forged user-agent arriving from a residential address is exactly what a forged user-agent arriving from a residential address looks like. Our verifier had no way to tell the difference, because there was none to see.
So every scan run by one of our own users pushed our published forgery rate up. The tool was measuring its own shadow.
How we found it #
Not by reading the code. A cluster of failed verifications arrived on one path in a few seconds wearing five different crawler names — a pattern no real crawler produces and no ordinary forger bothers with. Tracing the source address led back to our own network. The tell was in the timing, and it was only visible because the record keeps the path and the moment alongside the claim.
That is the same shape as an error we nearly published from access logs, where a 57% block rate turned out to be one machine wearing seven crawler names. Both were found by grouping on something other than the claimed identity — there, the source address; here, the path and the second. A user-agent is the one field in a request that carries no evidence about itself, so any count grouped only by user-agent is a count of what people typed.
The fix, and the part that matters #
Version 0.8.0 of the extension marks every probe it sends with a header identifying it as tooling. The receiving side reads that marker — but only to decide attribution, never to decide truth. A marked request is still recorded as a failed verification, because a header anyone can type must never be able to clear one. If the marker could clear a verdict, every real forger would send it and disappear from the count entirely.
Attribution is a claim about who; authentication is a claim about whether. Conflating them is how a measurement quietly becomes a bypass. The marker can only ever move a hit into a category that makes our own published number smaller, never larger.
The same asymmetry governs the opt-out header this site honours for its own traffic: a request can ask not to be counted, and that request is obeyed, but nothing a client sends can make it count as verified. Every self-declared header here is allowed to subtract and never to add. That is the only direction in which trusting a stranger's claim is safe.
What the record shows now #
Forged crawler identities #
7.2% of the requests that named themselves as a known crawler here were not that crawler. 1933 of 26767 checkable claims came from an address outside the range the operator publishes.
Forged rate by claimed identity — share of checkable claims from outside the operator’s published range
| Claimed to be | Claims | Verified | Forged | Unverifiable | Forged rate |
|---|---|---|---|---|---|
| Amazonbot | 20698 | 0 | 0 | 20698 | no feed |
| Meta-ExternalAgent | 9665 | 0 | 0 | 9665 | no feed |
| SemrushBot | 6985 | 0 | 0 | 6985 | no feed |
| Googlebot | 5640 | 5519 | 92 | 29 | 1.6% |
| PerplexityBot | 5234 | 5018 | 216 | — | 4.1% |
| AhrefsBot | 4591 | 0 | 0 | 4591 | no feed |
| GPTBot | 4299 | 4013 | 286 | — | 6.7% |
| ClaudeBot | 3799 | 3619 | 180 | — | 4.7% |
| Applebot | 3468 | 3235 | 233 | — | 6.7% |
| YandexBot | 2340 | 0 | 0 | 2340 | no feed |
| DataForSeoBot | 2229 | 0 | 0 | 2229 | no feed |
| OAI-SearchBot | 1805 | 1496 | 309 | — | 17.1% |
A user-agent is a claim, not an identity. Verified means the source IP sits inside a range the operator publishes — OpenAI, Anthropic, Google, Microsoft, Perplexity and Apple all publish one. Unverifiable is not forgery: some operators publish no range at all, so their requests can be neither confirmed nor accused, and they are excluded from the rate rather than counted against it. This is traffic to this site only, and this site is small — it is a floor on the problem, not a survey of the web.
Failed verifications are now split three ways. Requests that landed on report and tool endpoints are somebody checking a report we produced, not a crawler impersonating an operator on a page. Requests to actual content pages are the ones that mean what the headline implies. And hits recorded before the split shipped carry no path attribution at all — those stay in their own bucket, unassigned, rather than being guessed into whichever column flatters us.
The split matters more than we expected when we built it. Of every failed verification now on record, roughly half landed on real content pages — the ones the headline is about — while under three per cent came from tooling, and the remaining forty-five per cent predate the split and cannot be assigned to either. So the honest reading is a range, not a point: the content-only rate is roughly half the headline figure, and the true value sits somewhere between them until the unattributed bucket ages out. We publish both ends rather than the flattering one.
That third bucket is the largest, and it will stay in the record until it ages out. Backfilling it would have been one line of code and a lie.
The general trap: instruments that appear in their own data #
An instrument that touches the thing it measures will eventually count itself, and it is worth naming the ways rather than treating each one as a surprise.
This scanner fetches sites, so its own fetches land in those sites' logs. It records visits, so its own health checks would be visits. It publishes a corpus rate, and its own domain would be in that corpus. Each of those is handled by a deliberate exclusion: a header on every request this project makes to a site it operates, marking the traffic as its own and honoured as an opt-out; a set of addresses excluded from presence counting; and the fact that this scanner structurally cannot grade itself from inside, which we publish as a limit rather than hide.
The pattern in all of them is the same. You cannot remove the instrument from the world it measures, so the only options are to declare its footprint or to pretend it has none. Reporting a class you cannot resolve as its own bucket is the same discipline, applied to somebody else's opacity instead of your own.
Why publish this #
Because the alternative was to quietly restate a smaller number and let the old one keep circulating. A measurement business that corrects its own headline in public is making the only argument that actually distinguishes it: not that the number is impressive, but that you can check how it was reached.
And there is a narrower reason. The corrected figure is still large enough to make the original point, which is precisely why correcting it costs nothing but credibility gained. A finding that only survives on the strength of an error was never a finding.
Citing this figure. The numbers above are recomputed from the live record when this page loads, so link the anchor rather than freezing a copy: https://crawlcheck.io/blog/how-much-ai-crawler-traffic-is-forged#the-rate. Sentence form, as of 2026-10-07: 7.2% of requests claiming a verifiable AI-crawler identity, measured by CrawlCheck on its own traffic (1933 of 26767 adjudicable claims), came from outside the operator's published IP ranges; the most-forged name was Google-Extended at 100%.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
Can your own tooling corrupt your telemetry?
Yes. Our browser extension sends a crawler user-agent by design, and to this site's own telemetry that is indistinguishable from a stranger forging one.
How was the error found?
By reconciling the forged-claim record against what our own tools were known to send, rather than trusting the aggregate.
Why publish a mistake like this?
Because the number was published first. A corrected figure with the cause stated is worth more than a quiet edit.
Why does the path matter?
Because a request to a content page carrying a crawler name is the impersonation you care about, and a request to an API or report URL carrying one is usually tooling. Counting both as forgery inflates the rate and accuses the wrong people.
What is an unverifiable claim?
A request naming a crawler whose operator publishes no address ranges. It can be neither confirmed nor refuted, so it is excluded from every rate rather than counted as either.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.