CrawlCheck

Guides · 2026-10-03 · By · 0 views

How to verify GPTBot, ClaudeBot and PerplexityBot: tools compared

Where OpenAI, Anthropic, Perplexity and Google publish their crawler addresses, the tools that check traffic against them (CrawlCheck, Cloudflare, manual), and why unverifiable is its own answer.

To verify an AI crawler, match the request's IP against the operator's published list: OpenAI's gptbot.json, searchbot.json and chatgpt-user.json; Anthropic's ranges for ClaudeBot; Perplexity's perplexitybot.json; Google's reverse DNS or JSON files. CrawlCheck's free log verifier labels pasted log lines verified, spoofed or unverifiable, and Cloudflare's verified-bot signal checks requests at the edge on proxied sites.

Anyone can put “GPTBot” in a user-agent header. A log full of GPTBot requests, or of GPTBot 403s, says nothing until each request is checked against the addresses OpenAI actually publishes. The same holds for ClaudeBot, PerplexityBot and Googlebot. This guide lists the published sources, the tools that check against them, and what each one leaves out.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

Where each company publishes its crawler addresses #

OperatorCrawlersPublished IP list
OpenAIGPTBot, OAI-SearchBot, ChatGPT-Usergptbot.json, searchbot.json, chatgpt-user.json
AnthropicClaudeBot, Claude-SearchBot, Claude-UserPublished; linked from Anthropic’s crawler page
PerplexityPerplexityBot, Perplexity-Userperplexitybot.json, perplexity-user.json
GoogleGooglebot and other crawlersReverse then forward DNS to googlebot.com, google.com or googleusercontent.com, or the JSON range files in Google’s verification guide

Tools that verify crawler traffic #

ToolHow it verifiesWorks onPrice
CrawlCheck log verifierPaste raw access-log lines; each one naming a known crawler is checked against that operator’s published ranges and labelled verified, spoofed or unverifiableAny server’s logsFree
Cloudflare verified botsWeb Bot Auth signatures, published IP lists with a stable user-agent, or reverse DNS; usable in WAF custom rules and AI bot policiesSites proxied through CloudflareAvailable to Cloudflare customers
Manual checkDownload the JSON lists and match IPs yourself, or run reverse and forward DNS for GoogleAnythingFree; your time

Sources: Cloudflare verified bots, the operators’ pages above, and CrawlCheck’s free tools.

Why “unverifiable” is its own answer #

Not every crawler operator publishes addresses. A claim from a crawler with no published list cannot be proven real or fake, and treating it as either distorts the totals. CrawlCheck reports three outcomes and excludes unverifiable claims from every rate, so a share of “forged GPTBot traffic” is computed only over claims that could be checked. The method and the figures are in how much AI crawler traffic is forged.

Step by step: is this GPTBot request real? #

Checking one IP against a published list #

OpenAI’s and Perplexity’s files are JSON with a prefixes array of IP ranges in CIDR form, for example {"ipv4Prefix": "132.196.86.0/24"}, plus a creationTime. A request is from the crawler only if its address falls inside one of those ranges. Python’s standard library does the match:

import ipaddress, json, urllib.request
ranges = json.load(urllib.request.urlopen("https://openai.com/gptbot.json"))["prefixes"]
ip = ipaddress.ip_address("132.196.86.25")
print(any(ip in ipaddress.ip_network(r.get("ipv4Prefix") or r.get("ipv6Prefix")) for r in ranges))

Download the list again on a schedule rather than copying it once; the creationTime field changes as ranges are added and retired. Anthropic links its ranges from its crawler page; read that page for the current location and format.

Checking Googlebot with DNS #

Google’s method is two lookups, and both must agree. First a reverse lookup on the address from your log, which must return a name under googlebot.com, google.com or googleusercontent.com. Then a forward lookup on that name, which must return the original address. Google’s own example:

host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The forward lookup is the part people skip, and it is the part that matters: anyone who controls the reverse DNS for their own addresses can make them point at a googlebot.com name, but only Google can make that name resolve back.

Four mistakes that make crawler logs misleading #

What to do with each outcome #

OutcomeMeaningAction
Verified and servedThe real crawler read the pageNone
Verified and refused (403, 429, challenge)Your own rule is blocking the real crawlerFind the rule in the firewall, CDN or security plugin and allow the published ranges
SpoofedSomeone else using the crawler’s nameRefusing it is correct; do not loosen the rule
UnverifiableThe operator publishes no listDecide by behaviour (request rate, paths) rather than by name

Where CrawlCheck is weaker #

CrawlCheck checks the log lines you paste; it does not sit in front of your traffic, so it cannot block or rate-limit anything in real time. Cloudflare verifies every request at the edge as it arrives and can act on it, including through Web Bot Auth signatures. For sites on Cloudflare, use both: Cloudflare to enforce, CrawlCheck to audit logs from any host. Background: how to check if a GPTBot request in your logs is real.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

How do I verify that a request is really GPTBot?

Match the request's IP address against OpenAI's published list at openai.com/gptbot.json. A request that claims GPTBot from an address outside that list is spoofed. CrawlCheck's free log verifier does this for pasted log lines, and Cloudflare's verified-bot signal does it at the edge for proxied sites.

Where are ClaudeBot and PerplexityBot IP ranges published?

Anthropic publishes IP ranges for ClaudeBot, Claude-SearchBot and Claude-User, linked from its crawler support page. Perplexity publishes perplexitybot.json and perplexity-user.json on perplexity.com.

How does Google verify Googlebot?

Google recommends a reverse DNS lookup on the IP, confirming the name ends in googlebot.com, google.com or googleusercontent.com, then a forward lookup to the same IP; or matching against its published JSON range files.

What does Cloudflare's verified bot mean?

Cloudflare marks a bot as verified when it is transparent about who it is, confirmed by a Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS. The signal can be used in WAF custom rules on Cloudflare-proxied sites.

What does unverifiable mean in a crawler check?

The crawler's operator publishes no addresses to check against, so the claim cannot be proven real or fake. CrawlCheck reports it separately and excludes it from forgery rates.

Is fake GPTBot traffic common?

It happens often enough that a log of GPTBot 403s should be verified before you change any rule. CrawlCheck's measured share of forged crawler claims is published in its post on forged AI crawler traffic.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works