CrawlCheck

Findings · 2026-08-18 · By · 0 views

Why a self-identifying crawler header must never clear a check

A marker anyone can send must never be able to clear a check. It can label a request; it can never absolve one.

Our browser extension forges a user-agent on every probe. That is the job: to see what a site serves GPTBot, you have to ask as GPTBot. But to the receiving site’s telemetry that is indistinguishable from a stranger forging one — which is how we ended up counting our own users as forgers.

The fix was a header, x-cc-self, marking a request as tooling. The interesting decision was not adding it. It was refusing to let it mean anything.

The rule #

A marked request is still counted as failed verification. The verdict comes from the address check and nothing else. The marker changes only which bucket an already-failed check lands in: tooling rather than content.

The threat model is one line long. If the marker could clear a verification failure, every genuine forger would send it and disappear from the statistic entirely. We would have built a bypass and shipped it with documentation.

What the header may and may not do #

Written out as a permission table, because the temptation with a self-identifying marker is always to let it do one more useful thing, and every one of those things is the bypass.

EffectAllowed?Why
Change a verification verdict from failed to passedNeverA forger would send the header. The check exists to be unforgeable.
Move an already-failed check from the content bucket to the tooling bucketYesIt only ever makes our own published forgery rate smaller.
Exclude a request from the visitor record entirelyYesOur own fetches are not traffic to the site, and counting them would inflate every figure we publish.
Suppress a finding on the scanned siteNeverThe finding is about the site. Who asked has no bearing on what was served.
Raise a grade, unlock a paid check, or bypass a rate limitNeverAnything of value behind a header anyone can type is not behind anything.

The asymmetry that makes it safe #

Because the marker only ever moves a hit into a category that makes our own published number smaller, the incentive runs the right way. A self-identifying header that reduces your own reported problem is a defensible design. One that reduces your reported problem and hides other people’s is a backdoor with better branding.

Note which side of the line the third row sits on. Excluding our own fetches from the visitor record is not a favour to us: leaving them in would pad the traffic figures on our own dataset page. The header removes an inflation, and there is nothing a stranger could gain by sending it — the worst they can do is delete themselves from a private log nobody publishes.

The split has a denominator, and a gap we publish #

Separating tooling from content created a reporting problem we chose not to paper over. Hits recorded before the split shipped carry no path attribution at all: they cannot honestly be assigned to either side. So there are three counters, not two, and the third is the one most products would have quietly folded into the smaller number.

The rate itself is computed against verifiable claims only — requests naming a crawler whose operator publishes address ranges we can read. Counting the unverifiable ones in the denominator would treat no feed published as fine and understate every rate on the page. That exclusion is stated next to the figure rather than buried in a method note.

This is not a hypothetical about what a forged header could do. It is what arrives here. The rate below is read live as this page loads, so it moves with the traffic rather than with the publication date.

Forged crawler identities #

7.2% of the requests that named themselves as a known crawler here were not that crawler. 1933 of 26767 checkable claims came from an address outside the range the operator publishes.

Forged rate by claimed identity — share of checkable claims from outside the operator’s published range

Google-Extended100%Claude-SearchBot100%Perplexity-User89.6%Claude-User51.2%ChatGPT-User29.2%OAI-SearchBot17.1%Bingbot8.6%GPTBot6.7%Applebot6.7%ClaudeBot4.7%Google-Extended100%Claude-SearchBot100%Perplexity-User89.6%Claude-User51.2%ChatGPT-User29.2%OAI-SearchBot17.1%Bingbot8.6%GPTBot6.7%Applebot6.7%ClaudeBot4.7%
Claimed to beClaimsVerifiedForgedUnverifiableForged rate
Amazonbot206980020698no feed
Meta-ExternalAgent9665009665no feed
SemrushBot6985006985no feed
Googlebot5640551992291.6%
PerplexityBot52345018216—4.1%
AhrefsBot4591004591no feed
GPTBot42994013286—6.7%
ClaudeBot37993619180—4.7%
Applebot34683235233—6.7%
YandexBot2340002340no feed
DataForSeoBot2229002229no feed
OAI-SearchBot18051496309—17.1%

A user-agent is a claim, not an identity. Verified means the source IP sits inside a range the operator publishes — OpenAI, Anthropic, Google, Microsoft, Perplexity and Apple all publish one. Unverifiable is not forgery: some operators publish no range at all, so their requests can be neither confirmed nor accused, and they are excluded from the rate rather than counted against it. This is traffic to this site only, and this site is small — it is a floor on the problem, not a survey of the web.

Citing this figure. The numbers above are recomputed from the live record when this page loads, so link the anchor rather than freezing a copy: https://crawlcheck.io/blog/how-much-ai-crawler-traffic-is-forged#the-rate. Sentence form, as of 2026-10-07: 7.2% of requests claiming a verifiable AI-crawler identity, measured by CrawlCheck on its own traffic (1933 of 26767 adjudicable claims), came from outside the operator's published IP ranges; the most-forged name was Google-Extended at 100%.

What a site owner should conclude #

Nothing, and that is the point. If you operate a site and see x-cc-self on an inbound request, treat it as decoration: it tells you a request claims to be tooling, and a claim is not evidence. The only thing that ever settles who sent a request is whether the address sits inside a range the operator publishes, which is exactly what the free log verifier checks and what the crawler reference lists per operator.

We publish the header’s existence for the opposite reason to the usual one. Most self-identification exists so the sender can be let through. This one exists so our own probes can be excluded from figures we publish about other people’s traffic — and it is documented on the policy page so anyone can see what it does and confirm it does nothing else.

We proved it in both directions #

Before shipping, the same forged request was sent twice to a content page — once bare, once marked. The marked one landed in tooling, the bare one in content, and the failed-verification count went up both times. A test that only confirms the feature works is half a test; the half that matters is the one confirming it cannot be abused.

The same reasoning governs every other place a claim meets a check here. A user-agent is a claim; an IP range is evidence. A licence key proves what a customer paid for, never what a scanner found. And a fact nobody can check stays out of the score no matter how convenient it would be.

The rule generalises past headers. Any signal the subject of a measurement controls can be allowed to describe the measurement and never to decide it. A site can tell us it is a staging copy; that changes how a row is labelled, not whether the check ran. A licence can tell us who is paying; it changes what a customer may see, never what the scan found. The moment a controllable input can move a verdict, the verdict is measuring the input.

If you build self-identification into any measurement system, this is the line to hold: let it label, never let it absolve.

Applying the rule to your own telemetry #

If you record crawler visits, the check that matters is the one you probably do not have: does the source address fall inside a range the operator publishes? OpenAI, Anthropic, Perplexity, Google, Bing and Apple all publish theirs as JSON. A user-agent that fails that check is a claim, not a visit, and it belongs in a separate column. How much AI crawler traffic is forged reports what that column looks like across the sites we watch, and the crawler list names each operator's feed.

Once you have that column, resist the temptation this post is about. A header, a secret parameter, an allowlisted address, anything your own tooling sends to say this one is us, must only ever move a hit into a labelled bucket. The moment it can clear a verification failure, you have documented a bypass for everyone.

Where it lives in the report #

The telemetry section on every scan shows verified, spoofed and unverifiable hits as three numbers, never one, and the headline count is verified only. The glossary defines verified crawler and forged crawler identity; we counted our own users as forgers is the incident this header answers.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Should a self-identifying crawler header be trusted?

No. Anyone can send any header. A marker can label a request for convenience, but it must never be able to clear a check, because then the check measures nothing an attacker cannot fake.

What is the asymmetry that makes a header safe to use?

A header may only add suspicion, never remove it. Used that way a forged one gains the sender nothing, so there is no incentive to send it falsely.

How was the rule proved?

In both directions: a request carrying the header still had to pass every check, and a request without it was never penalised for the absence.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works