Guides · 2026-10-02 · By VSNARY | Emmanuel Orta · 0 views
How to verify Googlebot, Bingbot and AI crawlers by IP
A user-agent string is a claim anyone can type. Verification is reverse DNS with a forward confirm, or a match against the operator's published IP list. Which operators support which method, the file for each, and what to do with a crawler that publishes neither.
To verify Googlebot, run a reverse DNS lookup on the request's IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then run a forward lookup on that name and confirm it returns the same IP; or match the IP against Google's published JSON range files such as common-crawlers.json. Bingbot verifies the same way against search.msn.com or bingbot.json. OpenAI, Anthropic, Perplexity and Apple publish IP lists. A crawler whose operator publishes neither ranges nor a DNS convention cannot be verified, only reported as unverifiable.
Every crawler announces itself with a user-agent string, and every user-agent string can be copied. Scrapers routinely send Googlebot or GPTBot because sites tend to let those names through. Before you allow, block, rate-limit or count a crawler by name, the request has to be verified against something its operator controls. There are two such things, DNS and published IP lists, and operators support different ones. This guide covers both methods and lists what each major crawler publishes.
Method one: reverse DNS with a forward confirm #
Take the IP address from the log line and look up its PTR record. For a genuine Google crawler the name ends in googlebot.com, google.com or googleusercontent.com depending on the crawler type, for example crawl-66-249-66-1.googlebot.com. Then look up that hostname's A or AAAA record and confirm it returns the original IP. The forward step is the one that matters: anyone who controls an IP block can set its PTR record to say googlebot.com, but only Google can make googlebot.com names resolve to that IP.
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1Bing uses the same scheme with names ending in search.msn.com. DNS verification is exact but slow at scale, because each new IP costs two lookups, so cache the result per IP for a day or more.
Method two: the operator's published IP list #
Most operators now publish their crawler addresses as JSON lists of CIDR ranges. Matching an IP against a cached copy of the right list is fast and needs no DNS. The detail that trips people up is the word right: Google publishes separate files for its common crawlers, special-case crawlers, and user-triggered fetchers, and they are not interchangeable. In particular, user-triggered-fetchers.json covers fetches made from customers' Google Cloud projects, so an IP in that file proves the request came from Google's cloud, not that Google sent it.
| Crawler | Operator publishes | Where |
|---|---|---|
| Googlebot and common Google crawlers | DNS convention and IP list | googlebot.com / common-crawlers.json |
| Google user-triggered agents | IP list | user-triggered-agents.json, user-triggered-fetchers-google.json |
| Bingbot | DNS convention and IP list | search.msn.com / bing.com/toolbox/bingbot.json |
| GPTBot, OAI-SearchBot, ChatGPT-User | IP lists, one per agent | openai.com/gptbot.json, searchbot.json, chatgpt-user.json |
| ClaudeBot and Anthropic agents | IP list | claude.com/crawling/bots.json |
| PerplexityBot, Perplexity-User | IP lists | perplexity.ai/perplexitybot.json, perplexity-user.json |
| Applebot | IP list | search.developer.apple.com/applebot.json |
Refresh the lists daily. Operators add ranges without notice, and a stale list turns real crawler traffic into apparent forgeries.
When a crawler cannot be verified #
Some crawlers publish no ranges and no DNS convention. Bytespider is the most visible example. For those, a request carrying the name is not verifiable either way, and the honest classification is unverifiable, not genuine and not forged. Counting it as genuine inflates the crawler's apparent reach; counting it as forged accuses real traffic. We keep such crawlers in a separate bucket and exclude them from every verified rate, as described in the crawlers nobody can check.
How much named crawler traffic is forged #
Among requests that carry a verifiable crawler name, a meaningful share fail verification. The rate varies sharply by name, because forgers copy the names sites are most likely to allow. Our measured split, recomputed from live telemetry, is in how much AI crawler traffic is forged, and the GPTBot-specific walkthrough is in how to check whether a GPTBot request is real.
Signed requests are the next method #
IP verification breaks down for crawlers that run on shared cloud infrastructure, because the same addresses serve many tenants. The emerging fix is cryptographic: the crawler signs each request with a key it publishes, and the site verifies the signature. That scheme, Web Bot Auth, is covered in how signed crawler requests work. Until it is widespread, DNS and IP lists remain the only checks most crawlers support.
Doing it at scale #
For a busy site, verify in two tiers. First, keep a local copy of every operator's IP list, refreshed daily, and match each request's IP against the list for the name it claims; this is a CIDR lookup and costs microseconds. Only when the name has no list, or the IP is not in it, fall back to DNS, and cache the DNS result per IP for at least a day. Record three outcomes, not two: verified, forged, and unverifiable. Match each name to its own list. A request claiming GPTBot from an IP in OpenAI's ChatGPT-User list is an OpenAI request, but it is not GPTBot, and treating the lists as one pool hides the difference between a training crawler and a user-triggered fetch.
Watch IPv6. Several operators crawl over IPv6, and a matcher that only parses IPv4 ranges will mark every IPv6 request unverified. And keep the raw evidence: the IP, the claimed name, which list or DNS name verified it, and when the list was last refreshed. A verification result you cannot re-derive is a result you cannot defend when someone disputes a block.
Behind a CDN #
If the site sits behind a CDN or reverse proxy, the IP in the origin's logs is the proxy's, not the crawler's. Verification has to run on the client IP the proxy forwards, such as CF-Connecting-IP on Cloudflare, and only when the request actually came from the proxy; otherwise anyone can send that header directly to the origin and claim any address. Many CDNs offer their own verified-bot classification, which is a reasonable first filter but not a substitute for your own records.
What to do with the result #
Verify before you act on a name. Allow-list verified crawlers at the firewall by IP list rather than by user-agent, so a forged name gets the treatment an unknown client gets. Rate-limit unverifiable names like any other client. And when you report crawler traffic, report verified, forged and unverifiable separately; a single number mixing them measures how often a name was typed. The AI crawler list shows which method applies to each crawler we track.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
How do I verify that a request is really from Googlebot?
Reverse-DNS the IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then forward-resolve that name and confirm it returns the same IP. Or match the IP against Google's common-crawlers.json.
Why is the forward DNS lookup necessary?
Anyone who controls an IP block can set its reverse DNS to claim googlebot.com. Only Google can make googlebot.com names resolve to an IP, so the forward lookup is the proof.
How do I verify Bingbot?
The same way: reverse DNS should end in search.msn.com and forward-resolve to the original IP, or match the IP against Bing's published bingbot.json.
Do OpenAI and Anthropic publish crawler IP ranges?
Yes. OpenAI publishes separate JSON lists for GPTBot, OAI-SearchBot and ChatGPT-User, and Anthropic publishes a list at claude.com/crawling/bots.json.
Is an IP in user-triggered-fetchers.json proof that Google sent the request?
No. That file covers fetches from customers' Google Cloud projects, which anyone can run. It proves Google Cloud, not Google.
What if a crawler publishes no IP ranges?
Then a request carrying its name cannot be verified. Classify it as unverifiable and keep it out of verified and forged counts alike.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.