CrawlCheck

Guides · 2026-10-03 · By · 2 views

How to test robots.txt for GPTBot and ClaudeBot: 5 tools

Policy testers tell you whether your file allows a crawler; only a delivery check tells you whether the crawler gets the page. Five tools compared, the two-minute test, and the mistakes we see most.

To test robots.txt for GPTBot or ClaudeBot, run your live file through a checker that accepts any user-agent; CrawlCheck's free checker resolves it for 114 user-agents at once. Then check delivery: fetch the page as that crawler, because a firewall or CDN can refuse it even when robots.txt allows it. Google Search Console reports only Google's own fetch, and Google's open-source parser tests the file but not delivery.

Testing robots.txt for AI crawlers takes two checks, and most tools do only the first. The first is policy: given your file, is GPTBot (or ClaudeBot, or PerplexityBot) allowed to fetch a given URL? The second is delivery: when that crawler actually asks, does your server or CDN serve the page, or refuse it regardless of what robots.txt says? A file that allows GPTBot, behind a firewall that answers it with a 403, reads as “allowed” in every policy tester.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

Key points
  • A robots.txt test has two parts: policy (what the file allows) and delivery (what the crawler is actually served).
  • Under RFC 9309 a crawler named in its own group ignores the User-agent: * group entirely, so adding an AI crawler’s group can silently drop your Disallow rules.
  • A 4xx on robots.txt means “no rules”; a 5xx means “disallow everything” while it lasts.
  • Across CrawlCheck’s scans, robots.txt allows an answer engine’s crawler while the edge refuses it on 5.9% of sites.
  • CrawlCheck’s free checker resolves the live file for 114 user-agents at once; Google’s tools cover Google’s own crawlers.

What does testing robots.txt for GPTBot actually involve? #

Policy versus delivery: the two checks a robots.txt test needsCHECK 1 — POLICYRead the fileIs GPTBot allowed to fetch /pricing/under this file?User-agent: GPTBotAllow: /ANSWER: ALLOWEDCHECK 2 — DELIVERYFetch the pageAsk for /pricing/ as GPTBot andcompare with a browserbrowser 200 48 KBGPTBot 403 1 KBANSWER: REFUSEDSame site, same moment, opposite answers.A firewall or CDN decides before robots.txt is ever read, so a policy test alone reports “allowed”.
Most robots.txt testers stop at check 1. The crawler lives with the result of check 2.

Policy and delivery are decided in different places. The file lives on your server and is read by the crawler. Delivery is decided before that, by whatever answers the request first: a CDN, a firewall, a security plugin, a hosting provider’s bot filter. When those layers refuse a crawler, robots.txt is never consulted, and no policy tester can know, because a policy tester only reads the file.

The tools compared #

ToolTests any user-agent (GPTBot, ClaudeBot)?Checks delivery, not just the file?Cost
CrawlCheck robots.txt checkerYes: resolves the live file for 114 user-agents at once, group by group, most specific match firstYes: the full scan compares the policy with what each crawler identity was actually servedFree, no signup
Google Search Console robots.txt reportShows the robots.txt files Google found, when they were crawled, and parsing warnings or errorsNo; it reports Google’s own fetch of the fileFree, verified site owners
Google’s open-source robots.txt parserYes: you pass any user-agent and URLNo; it evaluates the file onlyFree; needs a developer to build and run
TechnicalSEO.com robots.txt tester (Max Prin)Tests whether a URL is blocked and howNoFree web tool
Cloudflare AI Crawl ControlTracks robots.txt health and which crawlers violate your directives, on a Cloudflare zoneSees requests that reach your zone; does not test from outsideIncluded on Cloudflare plans

Sources: Google’s robots.txt report help page, Google’s robotstxt library, TechnicalSEO.com, Cloudflare AI Crawl Control.

How does a crawler choose which rules apply to it? #

A robots.txt file is a list of groups. Each group starts with one or more User-agent lines and carries Allow and Disallow rules. RFC 9309 sets three rules for reading it that every AI crawler test should apply:

How a crawler chooses its group in robots.txtTHE FILEUser-agent: *Disallow: /cart/Disallow: /account/User-agent: GPTBotAllow: /blog/User-agent: ClaudeBotDisallow: /WHAT EACH CRAWLER OBEYSGPTBotonly its own group: /cart/ and /account/ are now allowedClaudeBotits own group: everything disallowedPerplexityBot (not named)the * group: /cart/ and /account/ disallowedA crawler named in its own group ignores the * group entirely. Rules are not inherited.
This is the most common robots.txt surprise we measure: adding a group to allow one AI crawler silently drops the Disallow rules written for everyone. Under RFC 9309, a crawler uses the group that names it, and the * group only when no group does.

The third rule causes most real-world damage. An owner adds User-agent: GPTBot with Allow: /blog/ to welcome OpenAI’s crawler, and in doing so removes every Disallow written under * for GPTBot: the cart, the account pages, the staging paths. Nothing errors and nothing looks different in a browser. The fix is to repeat the shared Disallow lines inside each named group. Our own robots.txt work hit this exact trap; the full explanation is in why naming a crawler drops your Disallow rules.

How to test robots.txt for GPTBot in two minutes #

  1. Open the checker. Go to robots.txt for AI crawlers and enter your domain, plus a path if you care about a page other than the homepage.
  2. Read the per-crawler result. Each AI crawler is listed as allowed or disallowed for that path, grouped by role: answer engines, search, training.
  3. Check the shadowed groups. The checker lists named groups that silently drop rules from the * group.
  4. Check the file itself. Run curl -sI https://example.com/robots.txt and confirm status 200 and content-type: text/plain. An HTML page at that address is not a robots.txt file, whatever its status.
  5. Run the full scan. From the homepage, scan the domain to see whether each crawler is actually served the page your file permits.

How often do these problems occur? #

robots.txt and access problems on scanned sites — live share of every CrawlCheck scan, read when this page loaded
robots.txt allows, the edge refuses5.9%
Named groups drop the * rules13.4%
An AI opt-out already set5%
robots.txt itself refused or erroring0.5%

A policy tester catches the second and third rows. Only a delivery check catches the first, and only a fetch of the file itself catches the fourth.

What do status codes on robots.txt do? #

The response to the robots.txt request itself changes everything that follows, and the rules are not intuitive:

What a crawler does with each robots.txt responseGET /robots.txt200 OKreads the rulesand caches themup to 24 hours3xx redirectfollows them,at least fivein a row4xx (403, 404)treats the file asabsent: may fetchanything5xx (500, 503)assumes completedisallow whileit lasts
From RFC 9309. A 403 on robots.txt means “no rules”, not “keep out”; a 503 means “keep out” for as long as it lasts. Neither is what most owners intend, so serve the file with status 200 and content type text/plain.

Two cases cause most trouble. A security layer that answers 403 to unknown clients on every path, including /robots.txt, tells a crawler there are no rules, and then refuses the pages anyway, so the crawler gets nothing and the owner’s file is never read. A server under load that answers 503 tells every crawler to stay away completely until it recovers. Detail and examples: robots.txt status codes: what 404, 403 and 503 tell a crawler.

Which AI crawlers should you test? #

Which AI crawler does whatTRAININGSEARCH / ANSWERSUSER-INITIATED FETCHOpenAIGPTBotOAI-SearchBotChatGPT-UserAnthropicClaudeBotClaude-SearchBotClaude-UserPerplexity—PerplexityBotPerplexity-User
Each name is a separate robots.txt decision. Blocking a training crawler does not remove a site from that company’s search answers; user-initiated fetchers act for a person, and operators treat robots.txt differently for them, so read each operator’s documentation.

Test each one you care about separately, because each is controlled separately. OpenAI documents GPTBot for training, OAI-SearchBot for ChatGPT search and ChatGPT-User for pages a person asks ChatGPT to open. Anthropic documents ClaudeBot, Claude-SearchBot and Claude-User in the same three roles. Perplexity documents PerplexityBot for its answers and Perplexity-User for user requests. Disallowing GPTBot does not remove a site from ChatGPT search; that is OAI-SearchBot. The full list, with each operator’s own documentation, is on the AI crawler list.

The mistakes we see most #

MistakeWhat happensHow to spot it
robots.txt allows the crawler, the firewall refuses itThe crawler never reads the page; policy testers say “allowed”Fetch the page as the crawler and as a browser and compare
A named group without the shared rulesDisallow lines meant for everyone stop applying to that crawlerThe checker’s shadowed-groups list
robots.txt answering 403 or 5034xx reads as no rules; 5xx as full disallowcurl -sI on the file
robots.txt served as HTMLNo directives can be parsed, so no rules applyThe content-type header
Blocking the training crawler to leave searchThe site stays in, or drops out of, the wrong productCheck each operator’s crawler roles
Editing the file and testing a cached copyThe test reads the new file; crawlers get the old oneFetch the bare URL and a cache-busted URL and compare

When should you test again? #

Re-test after any edit to robots.txt, after installing or updating an SEO or security plugin (either can rewrite the file or add bot rules), after moving host or CDN, and after switching on any bot-protection feature. Under RFC 9309 crawlers should not use a cached robots.txt for more than 24 hours, so a change can take up to a day to reach them, and a cached copy at your own edge can hold it back longer. Test the bare URL, not a cache-busted one, because the bare URL is what crawlers request.

Where CrawlCheck is weaker #

CrawlCheck sends each crawler’s published user-agent but cannot send from OpenAI’s or Anthropic’s own addresses, so a firewall rule keyed only to those IP ranges is reported as a difference between identities rather than confirmed. It reads the live file on your domain and does not let you test a draft before publishing; Google’s open-source parser does that, if you can build it. Search Console, unlike CrawlCheck, shows Google’s own fetch history for your file, so for Googlebot use it alongside.

Summary #

A policy tester answers “does my file allow GPTBot?”. Only a delivery check answers “does GPTBot actually get the page?”. Run both: the free robots.txt checker for the policy across every AI crawler, then a scan for delivery. More: do AI crawlers respect robots.txt?

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

How do I test robots.txt for GPTBot?

Enter your domain in a robots.txt checker that accepts any user-agent, such as CrawlCheck's free robots.txt checker, which resolves the live file for 114 user-agents including GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. Then confirm delivery with a scan, because a firewall can refuse GPTBot even when robots.txt allows it.

Does Google Search Console test robots.txt for GPTBot?

No. Search Console's robots.txt report shows the files Google found for your site, when they were crawled, and any warnings or errors. It reports Google's own fetch, not other companies' crawlers.

Why does my robots.txt allow GPTBot but ChatGPT still cannot read my site?

Usually a firewall, bot-management or CDN rule refuses the crawler before robots.txt matters, or the page serves a challenge. A delivery check that fetches as GPTBot shows the actual response.

If I add a group for GPTBot, does it still follow my User-agent: * rules?

No. Under RFC 9309 a crawler uses the group that names it and ignores the * group, so any Disallow rules you want GPTBot to obey must be repeated inside its own group.

Which rule wins when Allow and Disallow both match?

The most specific match, meaning the rule with the longest matching path. If an Allow and a Disallow match equally, RFC 9309 says the Allow should be used.

Does blocking GPTBot remove my site from ChatGPT search?

No. OpenAI documents GPTBot for training and OAI-SearchBot for ChatGPT search results. Each is controlled by its own robots.txt group.

What happens if robots.txt returns 403 or 503?

Under RFC 9309, a 4xx response means crawlers may treat the file as absent and fetch anything, and a 5xx means they must assume complete disallow while it lasts. Serve robots.txt with status 200 and content type text/plain.

Can I test robots.txt rules without publishing them?

Google's open-source robotstxt parser evaluates any file, user-agent and URL locally if you can build and run it. CrawlCheck's checker reads the live file on your domain.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects pages for training OpenAI's models. OAI-SearchBot fetches pages for ChatGPT search, and ChatGPT-User fetches a page when a person asks ChatGPT to open it. Each follows its own robots.txt group, so you can block GPTBot and still allow OAI-SearchBot. OpenAI says sites that disallow OAI-SearchBot are not shown in ChatGPT search answers, and that a robots.txt change takes about 24 hours to apply to search.

How do I check server logs for GPTBot requests?

Filter the access log for the user-agent string GPTBot, then check each request's IP address against the ranges OpenAI publishes at openai.com/gptbot.json (searchbot.json and chatgpt-user.json cover the other two). A request that names GPTBot from an address outside those ranges is not OpenAI. The status code on each line shows what the crawler actually received.

How do I change robots.txt on WordPress?

WordPress serves a generated robots.txt when there is no file in the site's root folder. Upload a real robots.txt to the root and it replaces the generated one; many SEO plugins also include a robots.txt editor. Then fetch the live /robots.txt to confirm the change, because a caching plugin or CDN can keep serving the old version.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works