Guides · 2026-10-03 · By VSNARY | Emmanuel Orta · 2 views
How to test robots.txt for GPTBot and ClaudeBot: 5 tools
Policy testers tell you whether your file allows a crawler; only a delivery check tells you whether the crawler gets the page. Five tools compared, the two-minute test, and the mistakes we see most.
To test robots.txt for GPTBot or ClaudeBot, run your live file through a checker that accepts any user-agent; CrawlCheck's free checker resolves it for 114 user-agents at once. Then check delivery: fetch the page as that crawler, because a firewall or CDN can refuse it even when robots.txt allows it. Google Search Console reports only Google's own fetch, and Google's open-source parser tests the file but not delivery.
Testing robots.txt for AI crawlers takes two checks, and most tools do only the first. The first is policy: given your file, is GPTBot (or ClaudeBot, or PerplexityBot) allowed to fetch a given URL? The second is delivery: when that crawler actually asks, does your server or CDN serve the page, or refuse it regardless of what robots.txt says? A file that allows GPTBot, behind a firewall that answers it with a 403, reads as “allowed” in every policy tester.
Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.
- A robots.txt test has two parts: policy (what the file allows) and delivery (what the crawler is actually served).
- Under RFC 9309 a crawler named in its own group ignores the
User-agent: *group entirely, so adding an AI crawler’s group can silently drop your Disallow rules. - A 4xx on robots.txt means “no rules”; a 5xx means “disallow everything” while it lasts.
- Across CrawlCheck’s scans, robots.txt allows an answer engine’s crawler while the edge refuses it on 5.9% of sites.
- CrawlCheck’s free checker resolves the live file for 114 user-agents at once; Google’s tools cover Google’s own crawlers.
What does testing robots.txt for GPTBot actually involve? #
Policy and delivery are decided in different places. The file lives on your server and is read by the crawler. Delivery is decided before that, by whatever answers the request first: a CDN, a firewall, a security plugin, a hosting provider’s bot filter. When those layers refuse a crawler, robots.txt is never consulted, and no policy tester can know, because a policy tester only reads the file.
The tools compared #
| Tool | Tests any user-agent (GPTBot, ClaudeBot)? | Checks delivery, not just the file? | Cost |
|---|---|---|---|
| CrawlCheck robots.txt checker | Yes: resolves the live file for 114 user-agents at once, group by group, most specific match first | Yes: the full scan compares the policy with what each crawler identity was actually served | Free, no signup |
| Google Search Console robots.txt report | Shows the robots.txt files Google found, when they were crawled, and parsing warnings or errors | No; it reports Google’s own fetch of the file | Free, verified site owners |
| Google’s open-source robots.txt parser | Yes: you pass any user-agent and URL | No; it evaluates the file only | Free; needs a developer to build and run |
| TechnicalSEO.com robots.txt tester (Max Prin) | Tests whether a URL is blocked and how | No | Free web tool |
| Cloudflare AI Crawl Control | Tracks robots.txt health and which crawlers violate your directives, on a Cloudflare zone | Sees requests that reach your zone; does not test from outside | Included on Cloudflare plans |
Sources: Google’s robots.txt report help page, Google’s robotstxt library, TechnicalSEO.com, Cloudflare AI Crawl Control.
How does a crawler choose which rules apply to it? #
A robots.txt file is a list of groups. Each group starts with one or more User-agent lines and carries Allow and Disallow rules. RFC 9309 sets three rules for reading it that every AI crawler test should apply:
- A crawler uses the group that names it. If several groups name it, their rules are combined. Only when no group names it does it fall back to the group for
*. - Within that group, the most specific rule wins, meaning the matching rule with the longest path. If an Allow and a Disallow match equally, the Allow should be used.
- Rules are not inherited. A crawler with its own group does not also obey the
*group.
The third rule causes most real-world damage. An owner adds User-agent: GPTBot with Allow: /blog/ to welcome OpenAI’s crawler, and in doing so removes every Disallow written under * for GPTBot: the cart, the account pages, the staging paths. Nothing errors and nothing looks different in a browser. The fix is to repeat the shared Disallow lines inside each named group. Our own robots.txt work hit this exact trap; the full explanation is in why naming a crawler drops your Disallow rules.
How to test robots.txt for GPTBot in two minutes #
- Open the checker. Go to robots.txt for AI crawlers and enter your domain, plus a path if you care about a page other than the homepage.
- Read the per-crawler result. Each AI crawler is listed as allowed or disallowed for that path, grouped by role: answer engines, search, training.
- Check the shadowed groups. The checker lists named groups that silently drop rules from the
*group. - Check the file itself. Run
curl -sI https://example.com/robots.txtand confirm status 200 andcontent-type: text/plain. An HTML page at that address is not a robots.txt file, whatever its status. - Run the full scan. From the homepage, scan the domain to see whether each crawler is actually served the page your file permits.
How often do these problems occur? #
A policy tester catches the second and third rows. Only a delivery check catches the first, and only a fetch of the file itself catches the fourth.
What do status codes on robots.txt do? #
The response to the robots.txt request itself changes everything that follows, and the rules are not intuitive:
Two cases cause most trouble. A security layer that answers 403 to unknown clients on every path, including /robots.txt, tells a crawler there are no rules, and then refuses the pages anyway, so the crawler gets nothing and the owner’s file is never read. A server under load that answers 503 tells every crawler to stay away completely until it recovers. Detail and examples: robots.txt status codes: what 404, 403 and 503 tell a crawler.
Which AI crawlers should you test? #
Test each one you care about separately, because each is controlled separately. OpenAI documents GPTBot for training, OAI-SearchBot for ChatGPT search and ChatGPT-User for pages a person asks ChatGPT to open. Anthropic documents ClaudeBot, Claude-SearchBot and Claude-User in the same three roles. Perplexity documents PerplexityBot for its answers and Perplexity-User for user requests. Disallowing GPTBot does not remove a site from ChatGPT search; that is OAI-SearchBot. The full list, with each operator’s own documentation, is on the AI crawler list.
The mistakes we see most #
| Mistake | What happens | How to spot it |
|---|---|---|
| robots.txt allows the crawler, the firewall refuses it | The crawler never reads the page; policy testers say “allowed” | Fetch the page as the crawler and as a browser and compare |
| A named group without the shared rules | Disallow lines meant for everyone stop applying to that crawler | The checker’s shadowed-groups list |
| robots.txt answering 403 or 503 | 4xx reads as no rules; 5xx as full disallow | curl -sI on the file |
| robots.txt served as HTML | No directives can be parsed, so no rules apply | The content-type header |
| Blocking the training crawler to leave search | The site stays in, or drops out of, the wrong product | Check each operator’s crawler roles |
| Editing the file and testing a cached copy | The test reads the new file; crawlers get the old one | Fetch the bare URL and a cache-busted URL and compare |
When should you test again? #
Re-test after any edit to robots.txt, after installing or updating an SEO or security plugin (either can rewrite the file or add bot rules), after moving host or CDN, and after switching on any bot-protection feature. Under RFC 9309 crawlers should not use a cached robots.txt for more than 24 hours, so a change can take up to a day to reach them, and a cached copy at your own edge can hold it back longer. Test the bare URL, not a cache-busted one, because the bare URL is what crawlers request.
Where CrawlCheck is weaker #
CrawlCheck sends each crawler’s published user-agent but cannot send from OpenAI’s or Anthropic’s own addresses, so a firewall rule keyed only to those IP ranges is reported as a difference between identities rather than confirmed. It reads the live file on your domain and does not let you test a draft before publishing; Google’s open-source parser does that, if you can build it. Search Console, unlike CrawlCheck, shows Google’s own fetch history for your file, so for Googlebot use it alongside.
Summary #
A policy tester answers “does my file allow GPTBot?”. Only a delivery check answers “does GPTBot actually get the page?”. Run both: the free robots.txt checker for the policy across every AI crawler, then a scan for delivery. More: do AI crawlers respect robots.txt?
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
How do I test robots.txt for GPTBot?
Enter your domain in a robots.txt checker that accepts any user-agent, such as CrawlCheck's free robots.txt checker, which resolves the live file for 114 user-agents including GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. Then confirm delivery with a scan, because a firewall can refuse GPTBot even when robots.txt allows it.
Does Google Search Console test robots.txt for GPTBot?
No. Search Console's robots.txt report shows the files Google found for your site, when they were crawled, and any warnings or errors. It reports Google's own fetch, not other companies' crawlers.
Why does my robots.txt allow GPTBot but ChatGPT still cannot read my site?
Usually a firewall, bot-management or CDN rule refuses the crawler before robots.txt matters, or the page serves a challenge. A delivery check that fetches as GPTBot shows the actual response.
If I add a group for GPTBot, does it still follow my User-agent: * rules?
No. Under RFC 9309 a crawler uses the group that names it and ignores the * group, so any Disallow rules you want GPTBot to obey must be repeated inside its own group.
Which rule wins when Allow and Disallow both match?
The most specific match, meaning the rule with the longest matching path. If an Allow and a Disallow match equally, RFC 9309 says the Allow should be used.
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI documents GPTBot for training and OAI-SearchBot for ChatGPT search results. Each is controlled by its own robots.txt group.
What happens if robots.txt returns 403 or 503?
Under RFC 9309, a 4xx response means crawlers may treat the file as absent and fetch anything, and a 5xx means they must assume complete disallow while it lasts. Serve robots.txt with status 200 and content type text/plain.
Can I test robots.txt rules without publishing them?
Google's open-source robotstxt parser evaluates any file, user-agent and URL locally if you can build and run it. CrawlCheck's checker reads the live file on your domain.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects pages for training OpenAI's models. OAI-SearchBot fetches pages for ChatGPT search, and ChatGPT-User fetches a page when a person asks ChatGPT to open it. Each follows its own robots.txt group, so you can block GPTBot and still allow OAI-SearchBot. OpenAI says sites that disallow OAI-SearchBot are not shown in ChatGPT search answers, and that a robots.txt change takes about 24 hours to apply to search.
How do I check server logs for GPTBot requests?
Filter the access log for the user-agent string GPTBot, then check each request's IP address against the ranges OpenAI publishes at openai.com/gptbot.json (searchbot.json and chatgpt-user.json cover the other two). A request that names GPTBot from an address outside those ranges is not OpenAI. The status code on each line shows what the crawler actually received.
How do I change robots.txt on WordPress?
WordPress serves a generated robots.txt when there is no file in the site's root folder. Upload a real robots.txt to the root and it replaces the generated one; many SEO plugins also include a robots.txt editor. Then fetch the live /robots.txt to confirm the change, because a caching plugin or CDN can keep serving the old version.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.