CrawlCheck

Guides · 2026-09-30 · By · 0 views

Cached robots.txt and llms.txt: stale, split, or blocked

Why one fetch cannot tell whether a machine file is healthy, the two-fetch method that can, the four findings it produces, and the purge order that clears them.

Machine files such as robots.txt, llms.txt and sitemap.xml are cached hardest by CDNs, so one fetch cannot tell whether the edge and the origin agree. CrawlCheck fetches each file bare and with a cache-buster and compares: different status codes are MACHINE_FILE_CACHE_STALE, same status but HTML versus not is MACHINE_FILE_CACHE_SPLIT, an edge HIT hiding an origin refusal is ORIGIN_BLOCKED_BEHIND_CACHE, and an old homepage copy that differs from a fresh one is STALE_CACHE_SERVED. Purge origin-side first, then the edge.

Share of all scans carrying each finding named aboveSTALE_CACHE_SERVED16.7%MACHINE_FILE_CACHE_STALE0.7%ORIGIN_BLOCKED_BEHIND_CACHE0.3%MACHINE_FILE_CACHE_SPLIT0.1%Share of all scans carrying eachfinding named aboveSTALE_CACHE_SERVED16.7%MACHINE_FILE_CACHE_STALE0.7%ORIGIN_BLOCKED_BEHIND_CACHE0.3%MACHINE_FILE_CACHE_SPLIT0.1%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

robots.txt, llms.txt, sitemap.xml and an entity map are small files that change rarely, which is exactly the profile a CDN caches hardest. That is fine until one of them changes, or until the origin starts refusing the crawler that asks for it, because from then on the edge keeps answering with whatever it last stored and nobody fetching the bare URL can tell. The scanner found this class of defect by being fooled by it: a bare request to a machine file returned a cached HTML interstitial at HTTP 200 while a cache-busted request to the same URL returned valid JSON. A single fetch cannot see that. This guide is the two-fetch method, the four findings it produces, and the purge order that actually clears it.

Why one fetch is not enough #

An edge cache is keyed on the URL. Ask for /robots.txt and you get the stored copy if there is one; the origin is not consulted. Ask for /robots.txt?cc=x7k2p and, unless the cache ignores query strings, the key misses and the request reaches the origin. Comparing the two answers tells you whether the edge and the origin currently agree about that file. They can disagree in status, in content type, or in bytes, and each disagreement means a different thing.

The four findings #

Status disagrees. The bare URL answers 200 and the origin answers 404, or the reverse. The edge is outliving a deleted or moved path, or a file that was just created has not reached the edge yet. Reported as MACHINE_FILE_CACHE_STALE (0.7% of scans). The scanner decides status first, deliberately: a 404 body is HTML by nature, so a status difference always looks like a content-type difference too, and an earlier version of the rule graded every stale path as the severe case.

Same status, different kind. Both answer 200, but one is HTML and the other is not. Some crawlers are getting a page where the file should be, usually a bot-challenge or maintenance interstitial that was cached under the file's URL. Reported as MACHINE_FILE_CACHE_SPLIT (0.1% of scans), and it is the case that caps a grade, because a crawler reading the cached copy is reading policy from an HTML page.

Edge hit, origin refuses. The bare URL answers 200 with cf-cache-status: HIT, and the cache-busted probe returns 403 or 503 from the origin. The file appears healthy from outside because the edge is serving a copy stored before the origin started refusing. The refusal is real and will surface the moment the cached copy expires. Reported as ORIGIN_BLOCKED_BEHIND_CACHE (0.3% of scans), with the blocker named where a signature identifies it.

Page served from an old copy. This one applies to the homepage rather than a machine file. The age header says how many seconds ago the stored copy was built. The scanner reads it, then fetches a fresh copy and compares: if the fresh copy differs or returns a different status, the visitor is getting an edit stuck behind the cache, and that is a real finding at higher severity; if the two are identical, the age is a note. Reported as STALE_CACHE_SERVED (16.7% of scans), with the age in hours and minutes.

Bare URLCache-bustedReadingCode
200404edge outlives a deleted pathMACHINE_FILE_CACHE_STALE
404200new file not at the edge yetMACHINE_FILE_CACHE_STALE
200 HTML200 JSON/textcached interstitial under the file's URLMACHINE_FILE_CACHE_SPLIT
200, HIT403 / 503origin refusing, edge hiding itORIGIN_BLOCKED_BEHIND_CACHE
homepage, age 9 h, bytes differfresh differsan edit stuck behind the cacheSTALE_CACHE_SERVED (severity up)
homepage, age 9 h, bytes equalfresh identicalold but correctSTALE_CACHE_SERVED (note)

Where the split actually comes from #

Almost every split case has the same origin: a bot-management or maintenance layer answered a request for the machine file with an HTML page, and that page was cacheable. The request that triggered it was usually from an unfamiliar address, which is exactly what a crawler is. The edge stored the interstitial under /llms.txt for the TTL the layer sent, and for that window every crawler asking for the file received a page telling it to enable JavaScript. The origin was fine the whole time; a cache-busted request proved it. The stale-status case is the mirror image: a file was deleted or moved at the origin and the edge kept answering 200 from its copy, which is harmless for a stylesheet and misleading for a policy file, because a robots.txt that was removed on purpose is still being obeyed.

Why the purge order matters #

A WordPress site on managed hosting behind a CDN has at least two caches in series: a server-side page cache such as Varnish, then the CDN edge. Purge the edge first and the next request pulls the still-stale copy out of Varnish and stores it at the edge again; the purge looked like it worked and changed nothing. The order that works is origin first, then edge: clear the server cache, confirm a direct request to the origin returns the new bytes, then purge the CDN. The post on a purge that returned 200 and evicted nothing is the case that established the rule, and a cache rule does not evict what is already cached covers the other trap: a new cache rule governs future stores, not the copies already sitting at the edge.

How to run the check yourself #

For each machine file, fetch the bare URL and the same URL with a random query string, and compare status, content type and size. Read the cf-cache-status or x-cache header on the bare response; a HIT with a mismatch is the edge, a MISS with a mismatch is something in front of the origin that keys on the query. Do it from a network that is not your own office, because a challenge page served to unfamiliar addresses is exactly the copy that gets cached under the file's URL. The report's machine-file section on crawlcheck.io runs this pair for every file it reads and shows both answers side by side; the sitemap guide and the robots.txt status guide cover what each file should return once the cache is out of the way.

Preventing it #

Give machine files a short edge TTL or exclude them from the page cache entirely; they are tiny and the origin cost is nothing. Never let a challenge or maintenance page be cacheable: a Cache-Control: no-store on every interstitial response is the single setting that prevents the split case. And when the origin changes what it serves to crawlers, whether a new firewall rule or a new hosting plan, purge the machine files explicitly rather than waiting for expiry, in the order above.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Why does the scanner fetch each machine file twice?

Once bare, as a crawler asks for it, and once with a random query string that misses the cache and reaches the origin. A single fetch cannot tell whether the edge and the origin agree; the pair can.

What is the difference between MACHINE_FILE_CACHE_STALE and MACHINE_FILE_CACHE_SPLIT?

Stale means the two fetches returned different status codes, usually an edge outliving a deleted path. Split means the same status but one answer is HTML and the other is not, which means some crawlers get a page where the file should be. Status is decided first because a 404 body is HTML by nature.

What does ORIGIN_BLOCKED_BEHIND_CACHE mean?

The bare URL returned 200 from the edge cache while a cache-busted request to the same path was refused by the origin with a 403 or 503. The file looks healthy only until the cached copy expires.

In what order should I purge caches?

Origin-side page cache first, then the CDN edge. Purging the edge first lets the next request re-store the still-stale server copy, so the purge changes nothing.

Is an old cached homepage a problem?

Only if it differs from the fresh copy. The scanner reads the age header, fetches a fresh copy and compares; an identical old copy is a note, a differing one is an edit stuck behind the cache and is reported at higher severity.

How do I stop a challenge page being cached under robots.txt?

Send Cache-Control: no-store on every challenge or maintenance response, and give machine files a short TTL or exclude them from the page cache.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works