Findings · 2026-10-06 · By VSNARY | Emmanuel Orta · 0 views
Our robots.txt check misread a full block as a broken rule
The shadowed-rules check compared rule text instead of what the rules cover, so sites that block a crawler outright were told their robots.txt was broken. Fixed on 6 October 2026.
CrawlCheck's shadowed-rules check flags robots.txt files where a named crawler's group leaves out rules from the wildcard group, because a crawler obeys only its own group. The check compared rule text instead of coverage, so a group with Disallow: /, which blocks every path, was reported as missing the wildcard's rules. 83% of sites blocking an AI crawler carried the false finding. It was fixed on 6 October 2026.
This is a correction. One of our robots.txt checks told sites that blocking a crawler completely had broken their rules for that crawler. It had not. The check was wrong, the fix shipped on 6 October 2026, and this post explains what happened.
What the check is for #
robots.txt groups do not add up. A crawler reads the one group that names it most specifically and ignores every other group, including the wildcard User-agent: *. So if a file disallows /admin/ for everyone and then adds a separate group for one crawler, that crawler no longer sees the /admin/ rule unless the new group repeats it. The user-agent groups guide covers the mechanism. It is a real and common mistake, and the check exists to catch it.
What went wrong #
The check decided whether a named group had dropped a rule by looking for the same rule text in that group. That works when a group lists paths one by one. It fails when a group blocks everything. The most common way to block a crawler is two lines:
User-agent: GPTBot Disallow: /
Disallow: / covers every path on the site, including /admin/. But the text / is not the text /admin/, so the check reported the rule as missing and told the site that GPTBot could now reach /admin/. It could not reach anything.
How we found it #
We found it while preparing the October dataset report. 83% of sites that block an AI crawler carried the finding, against 13% of sites that do not. A defect that is real and independent of blocking would not cluster like that. Following the flagged sites back to their robots.txt files showed that the blocking group itself was the trigger.
The fix #
The check now asks whether each wildcard rule is covered by any rule the named group declares, the way a crawler matches paths: a rule covers every path that starts with it, a * matches any run of characters, and a $ anchors the end. A group with Disallow: / covers everything, so it is never reported. A group that repeats a broader rule, such as /wp- in place of /wp-admin/, is not reported either.
| Named group | Wildcard group has | Before the fix | After the fix |
|---|---|---|---|
Disallow: / | Disallow: /admin/ | Reported as shadowed | Not reported: everything is blocked |
Disallow: /* | Disallow: /cart | Reported as shadowed | Not reported |
Disallow: /wp- | Disallow: /wp-admin/ | Reported as shadowed | Not reported: the broader rule covers it |
Allow: / only | Disallow: /admin/ | Reported | Still reported: the crawler can reach /admin/ |
| Repeats some rules, not all | Several Disallow rules | Reported | Still reported, for the rules it leaves out |
What still counts #
The real version of this mistake is untouched. A named group that allows everything, or repeats some wildcard rules and forgets others, still lets that crawler into paths the site meant to close, and it is still reported. On large sites we re-scanned after the fix, the finding still fires where a search crawler's group skips the wildcard's rules, which is exactly the case it exists for.
What changed for sites already scanned #
The fix shipped as a new rule version, so a site scanned before 6 October keeps its old reading until it is scanned again. When a site re-scans, the finding clears with the reason "cleared by a rule change" rather than appearing as a fix the site made, because the site did not change; our reading did. Certificates are re-measured automatically when the rule version moves.
If you block AI crawlers #
Blocking a crawler with Disallow: / in its own group was correct all along. If a CrawlCheck report told you otherwise, re-scan and the finding will clear. If you block some paths for a named crawler rather than all of them, make sure that group repeats every wildcard rule you still want applied; the robots.txt testing guide shows how to check a file from the crawler's side.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What is a shadowed robots.txt rule?
A rule in the User-agent: * group that a named crawler never sees, because the crawler obeys only the group that names it and that group does not repeat the rule.
Does Disallow: / in a named group shadow my other rules?
No. Disallow: / blocks every path for that crawler, including every path the wildcard group disallows. CrawlCheck reported it as shadowing before 6 October 2026; that was our error and it is fixed.
Why did my report change without my robots.txt changing?
The check was corrected and shipped as a new rule version. When a site re-scans, the finding clears marked as cleared by a rule change, not as a change to the site.
Which robots.txt files are still reported?
Files where a named crawler's group allows everything, or repeats only some of the wildcard group's Disallow rules, so that crawler can reach paths the site closed for everyone else.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.