CrawlCheck

Guides · 2026-10-02 · By · 0 views

robots.txt user-agent groups: why named bots skip your rules

A crawler obeys exactly one group in robots.txt, the most specific one that names it. Add a group for GPTBot with Allow: / and every Disallow under User-agent: * stops applying to GPTBot. How to see it, and how to write the file so it cannot happen.

Under RFC 9309 a crawler follows exactly one group in robots.txt: the group whose User-agent line matches it most specifically. If none names it, it follows User-agent: *. Groups are not combined, so adding User-agent: GPTBot with only Allow: / removes every Disallow rule in the * group for GPTBot. CrawlCheck raises ROBOTS_RULES_SHADOWED when a named group omits Disallow rules that the * group carries. The fix is to repeat the full * rule set inside every named group, or to list several agents on one group.

Share of all scans carrying each finding named aboveROBOTS_RULES_SHADOWED13.4%ROBOTS_IS_CATCHALL0.6%Share of all scans carrying eachfinding named aboveROBOTS_RULES_SHADOWED13.4%ROBOTS_IS_CATCHALL0.6%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Most robots.txt files that go wrong do not look wrong. Every line parses, every path is spelled correctly, and the homepage is reachable for every crawler you care about. The failure is structural: the file was edited by adding a block for a named crawler, usually an AI crawler, and that block quietly switched off the rules everyone assumed still applied. This guide explains the one rule behind it, how to check a file in under a minute, and how to write named groups that keep your exclusions intact.

How a crawler picks one robots.txt group and ignores the restTHE FILEUser-agent: *Disallow: /wp-admin/Disallow: /?s=Disallow: /feed/User-agent: GPTBotAllow: /WHAT GPTBOT OBEYSGPTBot group onlymost specific match* group ignored entirelyRESULT FOR GPTBOT/wp-admin/ open/?s= open/feed/ openRepeat the Disallows hereA named group does not add to *. It replaces it for that crawler. Repeat every * rule inside each named group.

The rule: one crawler, one group #

RFC 9309, the Robots Exclusion Protocol, says a crawler must obey the group whose User-agent line best matches its product token. If more than one group names the same crawler, those groups are merged into one. If no group names it, it obeys the User-agent: * group. What the standard does not do is combine a named group with the * group. The moment a group names a crawler, the * group no longer exists for that crawler. Google's robots.txt documentation describes the same behaviour for its crawlers.

So a file that reads User-agent: * followed by ten Disallow lines, and then User-agent: GPTBot followed by Allow: /, does not mean "everyone follows the ten rules and GPTBot is also welcome". It means GPTBot may fetch everything, including the ten paths you excluded. The same is true for any named crawler, including Googlebot and Bingbot.

Why it is so common now #

The pattern spread with AI crawler allowlists. Plugin snippets and blog advice recommend adding an explicit Allow: / block for GPTBot, ClaudeBot, PerplexityBot and a dozen others so that answer engines can reach the site. Each block is harmless on its own. Together they create a dozen more specific groups, each of which deletes the * rules for one crawler. Search, feed, login, cart and parameter URLs that were deliberately kept out of crawling become open to exactly the crawlers the edit was meant to welcome. We made this mistake ourselves: across sites we operate, the allowlist we deployed had shadowed every Disallow rule for every crawler it named, and our own scanner graded the files clean because each named crawler could reach the homepage.

The share of scanned sites where at least one named group drops rules that * carries is 13.4%, read from the live corpus each time this page is served.

How to check a file by hand #

Open the file at /robots.txt and walk it group by group. A group starts at one or more consecutive User-agent lines and runs until the next User-agent line that follows a rule. For each named group, list its Disallow lines and compare them with the * group. Every * Disallow that is missing from a named group is a rule that crawler does not follow.

File shapeWhat the named crawler obeysRules lost
* group onlythe * groupnone
* group + User-agent: GPTBot / Allow: /only Allow: /every * Disallow, for GPTBot
* group + named group that repeats every Disallowits own copy of the rulesnone
* declared twice, rules in the second blockboth * blocks, mergednone, but easy to misread
named group with its own different Disallowsonly its own Disallowsany * rule not repeated
/robots.txt returns the site's 404 page as 200nothing: no file existsall of them

Two traps make hand checks unreliable. First, a parser that closes a group at any unfamiliar line, such as Crawl-delay or Content-Signal, will attach the following rules to the wrong group; keep a group open across every directive until the next User-agent line that follows a rule. Second, groups that share a name merge, so a file that declares User-agent: * three times has one * group containing the union of all three. Count rules across every block with the same name before deciding anything is missing.

How to write named groups safely #

If you need a named group at all, give it the full rule set. The simplest correct shape lists every named crawler on consecutive User-agent lines above a single block of rules, which is legal and makes them one group:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /feed/

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: PerplexityBot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /?s=
Disallow: /feed/
Allow: /

Ask first whether you need the named group. A crawler with no group of its own already follows *, so a site that wants AI crawlers to have the same access as everyone else needs no AI section at all. Named groups are for treating a crawler differently: a stricter rule set, a Crawl-delay, or a full block. If a plugin or theme generates the file, patch the generator rather than the output, because the next save will regenerate the shadowed version.

What a shadowed group costs #

The cost is crawling and possibly indexing of URLs you excluded on purpose: internal search results, feeds, faceted parameters, login and account pages, staging paths. For search crawlers that means crawl budget spent on duplicates and thin pages, and sometimes those pages appearing in results. For AI crawlers it means the excluded pages can be fetched and used, which is usually the opposite of the intent behind the exclusion. None of it shows up as an error anywhere. Search Console will not warn you, because the file is valid.

Allow, Disallow and the longest match #

Inside the group a crawler has chosen, a second rule decides between conflicting lines: the most specific path wins, measured by the length of the matching rule, and when an Allow and a Disallow match with equal length, RFC 9309 says the Allow wins. That is why Allow: /wp-admin/admin-ajax.php can sit beside Disallow: /wp-admin/ and still open the one file. It is also why the order of lines inside a group does not matter to Google. It does matter to some older parsers, including Python's standard urllib.robotparser, which applies the first matching rule and does not understand * or $ wildcards. A file tested only with that parser can look broken when it is correct, or correct when it is broken. Test with a parser that implements RFC 9309, and put each group's specific Allow exceptions first, then its Disallow lines, then any blanket Allow: / last, so every parser reaches the same answer.

Wildcards follow the same longest-match rule. Disallow: /*?replytocom= blocks every URL carrying that parameter, and Disallow: /*.pdf$ blocks URLs ending in .pdf. When a named group copies the * rules, copy the wildcard rules too; they are the ones most often dropped, because they look like noise next to the readable paths.

What CrawlCheck reports #

The scanner parses every group with RFC 9309 merging, then compares each named group's Disallow set with the merged * set. ROBOTS_RULES_SHADOWED names how many groups and agents are affected and the largest number of rules any one of them drops, and it says so explicitly when a search crawler is among them, because those bypassed paths can enter the index. A file that answers with the site's catch-all page instead of a robots file is ROBOTS_IS_CATCHALL, currently 0.6% of scans; crawlers treat that as no file and allow everything. The broader question of which AI crawlers honour robots.txt at all is in the robots.txt compliance guide, and blocking one crawler correctly is covered in how to block AI crawlers. The robots.txt checker runs the group comparison on any domain.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Does a crawler combine its own robots.txt group with User-agent: *?

No. RFC 9309 says a crawler follows only the group that best matches its name. If a group names it, the * group does not apply to it at all.

What does Allow: / in a named group do?

It gives that crawler permission to fetch everything, and because the named group replaces the * group for that crawler, every Disallow rule in the * group stops applying to it.

How do I keep my Disallow rules when I add an AI crawler group?

Repeat every Disallow and Allow line from the * group inside the named group, or list several crawlers on consecutive User-agent lines above one shared block of rules.

Do I need a named group to allow AI crawlers?

No. A crawler with no group of its own already follows User-agent: *. Named groups are only needed to treat a crawler differently from everyone else.

What if User-agent: * appears more than once?

Groups that name the same agent merge into one, so the rules from every * block apply together. Count rules across all of them before concluding any are missing.

Will Search Console warn me about a shadowed group?

No. The file is syntactically valid, so robots.txt testers report no error. Only a comparison of each named group against the * group shows the lost rules.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works