Guides · 2026-10-03 · By VSNARY | Emmanuel Orta · 0 views
How to verify GPTBot, ClaudeBot and PerplexityBot: tools compared
Where OpenAI, Anthropic, Perplexity and Google publish their crawler addresses, the tools that check traffic against them (CrawlCheck, Cloudflare, manual), and why unverifiable is its own answer.
To verify an AI crawler, match the request's IP against the operator's published list: OpenAI's gptbot.json, searchbot.json and chatgpt-user.json; Anthropic's ranges for ClaudeBot; Perplexity's perplexitybot.json; Google's reverse DNS or JSON files. CrawlCheck's free log verifier labels pasted log lines verified, spoofed or unverifiable, and Cloudflare's verified-bot signal checks requests at the edge on proxied sites.
Anyone can put “GPTBot” in a user-agent header. A log full of GPTBot requests, or of GPTBot 403s, says nothing until each request is checked against the addresses OpenAI actually publishes. The same holds for ClaudeBot, PerplexityBot and Googlebot. This guide lists the published sources, the tools that check against them, and what each one leaves out.
Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.
Where each company publishes its crawler addresses #
| Operator | Crawlers | Published IP list |
|---|---|---|
| OpenAI | GPTBot, OAI-SearchBot, ChatGPT-User | gptbot.json, searchbot.json, chatgpt-user.json |
| Anthropic | ClaudeBot, Claude-SearchBot, Claude-User | Published; linked from Anthropic’s crawler page |
| Perplexity | PerplexityBot, Perplexity-User | perplexitybot.json, perplexity-user.json |
| Googlebot and other crawlers | Reverse then forward DNS to googlebot.com, google.com or googleusercontent.com, or the JSON range files in Google’s verification guide |
Tools that verify crawler traffic #
| Tool | How it verifies | Works on | Price |
|---|---|---|---|
| CrawlCheck log verifier | Paste raw access-log lines; each one naming a known crawler is checked against that operator’s published ranges and labelled verified, spoofed or unverifiable | Any server’s logs | Free |
| Cloudflare verified bots | Web Bot Auth signatures, published IP lists with a stable user-agent, or reverse DNS; usable in WAF custom rules and AI bot policies | Sites proxied through Cloudflare | Available to Cloudflare customers |
| Manual check | Download the JSON lists and match IPs yourself, or run reverse and forward DNS for Google | Anything | Free; your time |
Sources: Cloudflare verified bots, the operators’ pages above, and CrawlCheck’s free tools.
Why “unverifiable” is its own answer #
Not every crawler operator publishes addresses. A claim from a crawler with no published list cannot be proven real or fake, and treating it as either distorts the totals. CrawlCheck reports three outcomes and excludes unverifiable claims from every rate, so a share of “forged GPTBot traffic” is computed only over claims that could be checked. The method and the figures are in how much AI crawler traffic is forged.
Step by step: is this GPTBot request real? #
- Copy the log lines that name GPTBot (or any AI crawler) from your access log.
- Paste them into the log check on CrawlCheck’s free tools. Each line returns verified, spoofed or unverifiable, with the range it matched.
- If most “GPTBot” 403s are spoofed, your host is blocking an impostor, not OpenAI; leave the rule. If verified requests are refused, fix the rule that refuses them.
- On Cloudflare, use the verified-bot signal in rules rather than matching the user-agent string.
Checking one IP against a published list #
OpenAI’s and Perplexity’s files are JSON with a prefixes array of IP ranges in CIDR form, for example {"ipv4Prefix": "132.196.86.0/24"}, plus a creationTime. A request is from the crawler only if its address falls inside one of those ranges. Python’s standard library does the match:
import ipaddress, json, urllib.request
ranges = json.load(urllib.request.urlopen("https://openai.com/gptbot.json"))["prefixes"]
ip = ipaddress.ip_address("132.196.86.25")
print(any(ip in ipaddress.ip_network(r.get("ipv4Prefix") or r.get("ipv6Prefix")) for r in ranges))Download the list again on a schedule rather than copying it once; the creationTime field changes as ranges are added and retired. Anthropic links its ranges from its crawler page; read that page for the current location and format.
Checking Googlebot with DNS #
Google’s method is two lookups, and both must agree. First a reverse lookup on the address from your log, which must return a name under googlebot.com, google.com or googleusercontent.com. Then a forward lookup on that name, which must return the original address. Google’s own example:
host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1The forward lookup is the part people skip, and it is the part that matters: anyone who controls the reverse DNS for their own addresses can make them point at a googlebot.com name, but only Google can make that name resolve back.
Four mistakes that make crawler logs misleading #
- Blocking by user-agent string alone. It refuses the real crawler and anyone impersonating it alike, and the impersonators simply change the string.
- Treating unverifiable as fake. A crawler whose operator publishes no list cannot be proven either way; counting it as forged inflates the forged share.
- A range list copied once. Ranges change; a list from last year will mark new, genuine addresses as spoofed.
- Ignoring IPv6. If your server logs IPv6 addresses, check them against the IPv6 entries, not only the IPv4 ones.
What to do with each outcome #
| Outcome | Meaning | Action |
|---|---|---|
| Verified and served | The real crawler read the page | None |
| Verified and refused (403, 429, challenge) | Your own rule is blocking the real crawler | Find the rule in the firewall, CDN or security plugin and allow the published ranges |
| Spoofed | Someone else using the crawler’s name | Refusing it is correct; do not loosen the rule |
| Unverifiable | The operator publishes no list | Decide by behaviour (request rate, paths) rather than by name |
Where CrawlCheck is weaker #
CrawlCheck checks the log lines you paste; it does not sit in front of your traffic, so it cannot block or rate-limit anything in real time. Cloudflare verifies every request at the edge as it arrives and can act on it, including through Web Bot Auth signatures. For sites on Cloudflare, use both: Cloudflare to enforce, CrawlCheck to audit logs from any host. Background: how to check if a GPTBot request in your logs is real.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
How do I verify that a request is really GPTBot?
Match the request's IP address against OpenAI's published list at openai.com/gptbot.json. A request that claims GPTBot from an address outside that list is spoofed. CrawlCheck's free log verifier does this for pasted log lines, and Cloudflare's verified-bot signal does it at the edge for proxied sites.
Where are ClaudeBot and PerplexityBot IP ranges published?
Anthropic publishes IP ranges for ClaudeBot, Claude-SearchBot and Claude-User, linked from its crawler support page. Perplexity publishes perplexitybot.json and perplexity-user.json on perplexity.com.
How does Google verify Googlebot?
Google recommends a reverse DNS lookup on the IP, confirming the name ends in googlebot.com, google.com or googleusercontent.com, then a forward lookup to the same IP; or matching against its published JSON range files.
What does Cloudflare's verified bot mean?
Cloudflare marks a bot as verified when it is transparent about who it is, confirmed by a Web Bot Auth signature, a published IP list with a stable user-agent, or reverse DNS. The signal can be used in WAF custom rules on Cloudflare-proxied sites.
What does unverifiable mean in a crawler check?
The crawler's operator publishes no addresses to check against, so the claim cannot be proven real or fake. CrawlCheck reports it separately and excludes it from forgery rates.
Is fake GPTBot traffic common?
It happens often enough that a log of GPTBot 403s should be verified before you change any rule. CrawlCheck's measured share of forged crawler claims is published in its post on forged AI crawler traffic.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.