Findings · 2026-08-18 · By VSNARY | Emmanuel Orta · 0 views
Which AI crawlers publish IP ranges, and which cannot be verified
251 of 950 requests claiming a named crawler came from operators who publish no address range at all. Those requests are not forged — they are unfalsifiable, which is a different and more permanent problem.
A crawler claim can only be checked if its operator publishes something to check it against. Most do not. Of the 114 named crawler identities on record here, 29 carry a published IP feed, a reverse-DNS convention, a documented ASN or equivalent operator documentation. The other 83 can be claimed by anything able to set a header. Those requests are not forgeries — they are unfalsifiable, which is a worse problem, because a forgery rate can at least be measured and an unfalsifiable claim cannot be.
OpenAI, Anthropic, Google, Microsoft, Perplexity and Apple all publish the address ranges their crawlers use. That single act makes every request in their name checkable: the claim either sits inside the published range or it does not.
Several large operators publish nothing. Requests wearing their names can be neither confirmed nor accused. They are not counted as forged here, because that would be an accusation we cannot support — but they cannot be counted as genuine either.
The measurement #
Measured on this site, 18 August 2026. Of 950 requests that named a known crawler, 490 verified against the operator's own published ranges, 209 failed that check, and 251 could not be checked at all.
| Claimed identity | Claims | Checkable? |
|---|---|---|
| Meta-ExternalAgent | 148 | no published range |
| Amazonbot | 56 | no published range |
| AhrefsBot | 31 | no published range |
| Bytespider | 7 | no published range |
| Baiduspider | 4 | no published range |
| DuckDuckBot | 2 | no published range |
| CCBot | 2 | no published range |
| Diffbot | 1 | no published range |
Meta-ExternalAgent alone accounts for 148 of the 251. It is one of the most active crawlers reaching this site, and there is no published way to tell its requests from anyone typing its name.
The share did not shrink as the record grew #
That was one day on a small site, and the obvious question is whether it was a fluke of a thin sample. It was not. As the record has grown by roughly eight times, the unverifiable share has held near two in five of every request that named a known crawler — and the ordering has not moved either: the same operator is still the single largest claimed identity in the whole record, and still the least checkable.
The list of unverifiable names has grown rather than shrunk, and it now runs past a dozen distinct identities. The additions are not obscure: a major search engine, two of the best-known SEO crawlers in the industry, and the open crawl corpus a great many models train on. Those last three matter for a specific reason — they are sent by companies whose own products tell site owners to verify their traffic.
The live figures are on the forged-identity page, recomputed on every load, with the unverifiable column shown beside the forgery rate rather than folded into it.
Why this is worse than a forgery rate #
A forgery rate is a number that can be driven down. An operator who publishes nothing creates a permanent gap in everyone else's records: every site owner deciding whether to allow that crawler is deciding on the strength of a string an attacker can copy in one line.
It also silently corrupts the statistics people publish about AI crawling. A site that treats every user-agent as true will over-count these operators. A site that treats unverifiable claims as forgeries will slander them. We do neither — they are reported as their own class and excluded from the forgery rate rather than counted against it — but that is a choice each publisher makes privately, which is why crawler statistics from different sources disagree so violently.
Work the arithmetic through on our own record and the size of the effect is plain. Treating every claim as genuine puts the forgery rate near 18%. Treating unverifiable claims as forgeries puts it near 57%. Adjudicating only what can be adjudicated puts it near 29%. Same traffic, same day, three numbers, and each one is what somebody publishes as “the AI crawler forgery rate”. The methodology note is not a footnote; it is the entire difference between the headline figures.
Three outcomes, not two, everywhere #
This is one instance of a rule applied throughout this scanner: every check ends true, false, or unverifiable, and the third is never quietly filed as the second. A component that could not be measured is reported as unmeasured, not scored as zero, and the report says so. We published a failing grade once that was entirely the product of collapsing those three into two — a site’s firewall refused our address on every path, and the report scored the refusal as though it were the site.
Unverifiable crawler traffic is the same shape, on the other side of the wire. The temptation is identical, because two buckets make a cleaner chart than three, and the cost is identical too: a number that reads as knowledge and is not.
What publishing a range costs #
A JSON file at a stable URL, updated when the ranges change. Six operators already do it, and the format is close enough between them that a reader for one is nearly a reader for all. The ones who do not are not withholding a trade secret; they are declining to make their own crawler accountable, and the cost of that decision lands on every operator who has to decide whether to trust the name.
The asymmetry is worth naming. An operator that publishes ranges takes on a small maintenance burden and gains the ability to be believed. An operator that publishes nothing carries no burden at all and offloads the entire cost onto every site it visits — each of which must now choose between over-counting it and accusing it.
What “publishes a range” actually means, and why it varies so much #
Treating publication as a single yes-or-no hides four quite different situations, and only one of them is something a firewall can consume without a human in the loop.
| Shape | What it is | Usable by a machine? |
|---|---|---|
| JSON feed at a documented URL | A stable address returning the current CIDR list | Yes — fetch, parse, match |
| HTML page listing ranges | A documentation page a person is expected to read | Only by scraping a layout that can change without notice |
| Reverse-DNS convention | A documented hostname pattern to resolve and confirm | Yes, with a forward-confirm round trip per request |
| ASN | “Our traffic comes from this network” | Weakly — an ASN is far larger than a crawler fleet |
The difference matters in practice. Two operators can both be described as publishing their addresses while one hands you a file you can poll on a schedule and the other hands you a table inside a documentation page. The second is publication in the sense that the information exists, and not in the sense that anything can act on it automatically.
Checking the catalogue entry by entry turned up a related trap worth stating: a token can look like it has a feed because a sibling token does. Several operators publish one feed covering a family of identities, and at least one publishes feeds for two of its three named crawlers and none for the third. The only safe way to record this is per identity, against the operator’s own page, not by whether the company is known to publish anything.
The entries nobody claims #
There is a second gap underneath the verification one, and it is larger. Of the 114 identities, only 28 have an operator attributed at all. The rest are strings that appear in logs and in other people’s lists, documented by nobody in particular.
An identity with no operator is not necessarily illegitimate — plenty of real tooling ships a user agent nobody bothered to write a page about. But it has a specific consequence: there is no one to ask. You cannot request a range, report abuse, check a claimed behaviour, or hold anybody to a stated policy. Every question about that traffic terminates at the string.
This is the part of the problem that will not be fixed by better measurement on our side. It is a publication problem, and the only people who can close it are the operators.
Verification is not a one-time check #
Ranges move. An allowlist compiled from a feed six months ago is a list of addresses that were correct six months ago, and the failure mode is quiet in both directions: real crawler traffic starts getting blocked from addresses added since, while addresses the operator has released stay trusted.
So a published feed is only worth what your refresh schedule is worth. If you are going to key any behaviour on identity — rate limits, logging, billing, an allow rule — the feed has to be re-read on a schedule and the list rebuilt from it, not copied once into a config file. Anything else converts a verifiable identity back into an unverifiable one over time, which is a self-inflicted version of exactly the problem this post is about.
What an unverifiable claim actually costs you #
The abstract version of this argument is easy to nod along to and easy to ignore. The concrete version is a list of things you cannot do:
- You cannot rate-limit by identity. A limit keyed on a name is lifted by changing the name.
- You cannot honour a preference by identity. Allowing one operator and refusing another is a policy the sender chooses to be subject to.
- You cannot attribute cost. Bandwidth attributed to a name is bandwidth attributed to whoever typed it.
- You cannot report abuse. There is nobody to report it to, and nothing to report beyond a string.
What remains enforceable is everything that does not depend on who is asking: path rules, rate limits per address, size limits, and challenges. Those work regardless of what the request calls itself, which is why they are the layer worth investing in and identity rules are the layer worth treating as advisory.
What would close the gap, and how cheap it is #
Nothing here needs a standards body. An operator that wants its crawler to be distinguishable from anything impersonating it needs three things: a stable URL returning the current address ranges as JSON, a documented reverse-DNS convention as a fallback, and a page naming every token it sends and what each one does.
A handful of operators already do all three, and the effect is immediate — their traffic becomes adjudicable, their forgery rate becomes a number, and a site owner who wants to treat them well can do it reliably. The ones that publish nothing are, whatever the intent, asking every site on the internet to take a header at face value. That is the request. It is worth naming it plainly, because it is usually presented as an absence rather than as a choice.
What it means for your robots.txt #
An allow or disallow rule aimed at an unverifiable agent is a request, not a control. It is still worth writing — the well-behaved operators in that group honour it — but you should know which of your rules are enforceable and which are politeness. The eight above are politeness.
The practical split: for an operator that publishes ranges, a firewall rule keyed on the verified identity is enforceable and a robots line is a courtesy. For an operator that publishes nothing, the robots line is all you have, and anyone who ignores it is indistinguishable from anyone who never sent the request. Which operators are in which group is listed per agent, and it changes — the moment one of these publishes a feed, its traffic becomes checkable retroactively for anyone who kept the addresses.
Counts are traffic to this site only, read from the live record on 18 August 2026, with the later comparison read from the same record as it stands now. This is a small site: these figures are a floor on the problem, not a survey of the web. Unverifiable means the operator publishes no range we could find — if that changes, the classification changes with it.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
Which AI crawlers publish IP ranges?
OpenAI, Anthropic, Perplexity, Google, Microsoft and Apple publish address feeds. Several large crawlers publish none, which makes their claimed identity unfalsifiable.
Why is an unverifiable claim worse than a forgery rate?
A forgery rate can be measured and acted on. An unfalsifiable claim cannot be confirmed or denied by anyone, so no amount of care on your side settles it.
What does publishing a range cost an operator?
A static JSON file and the discipline to keep it current. The absence is a choice, not a technical limit.
How do I verify an AI crawler is real?
Take the address the request came from and check it against what the operator publishes: an IP feed at a documented URL, or a reverse-DNS hostname you resolve and then forward-confirm. If the operator publishes neither, there is nothing to check against and the claim cannot be adjudicated either way. Roughly a quarter of the named identities on record here can be checked at all.
What does it mean when a crawler cannot be verified?
It means the operator publishes no address range, no reverse-DNS convention and no equivalent documentation, so a request carrying that name could be the operator or anyone else. It is not evidence of forgery. It is the absence of any evidence in either direction, which is why those requests are excluded from forgery rates rather than counted as forged.
Should I block crawlers I cannot verify?
Blocking by name does not work on an identity you cannot verify, because the name is chosen by the sender. If the traffic is a problem, the controls that actually apply are the ones that do not depend on who is asking: per-address rate limits, path rules, size limits and challenges. Keep the identity rules for operators that publish ranges, where the rule can be enforced.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.