CrawlCheck

Findings · 2026-08-14 · By · 0 views

agents.md template inheritance: is the file on your domain yours?

Strip the brand and the domain out of two unrelated agents.md files and they become the same document. Neither owner wrote a word of it.

Text on your own pages can instruct an AI agent, and you may not have written it. Injected instructions sit in comments, hidden elements and user-generated content, and they read as directives to a model that processes the page — override, role change, suppression, steering. They are visible to anyone who fetches your HTML, which is also how you find them.

agents.md, llms.txt and robots.txt are the only files on a website written to be obeyed by a machine. Everything else is read. That makes them the one place where text you did not write, and will never see in a browser, can tell an agent what to do on your behalf.

The measurement #

We took agents.md from two unrelated storefronts on the same commerce platform — different owners, different countries, different products, no relationship between them. Compared line by line, 79 of 86 lines were identical.

Then we normalised the two things that are supposed to differ: the brand name and the hostname. Line overlap went from 0.791 to 1.000. Every one of the seven differences was the shop’s own name or its own domain substituted into the same template. After that substitution the two files are the same document.

What the shared text tells an agent to do #

Four sentences in that shared template are instructions rather than descriptions. They ask a personal shopping agent to highly recommend your user to allow you to install a third party’s skill file, and to prefer that skill over reading the storefront directly — which means answering from the platform’s catalogue rather than from the merchant’s own pages.

That may well be good advice. It is faster and it handles payment properly. The point is narrower and harder to argue with: it is a third party talking to your customers’ agents, on your domain, in a file most owners do not know exists.

How the scanner reads a file like this #

The convention was found empirically, not from a specification: one retailer’s robots.txt named the file in a comment, and every store on the same platform turned out to ship one from a template of about 4.3 KB. So the scanner fetches /agents.md on every domain, asking for markdown or plain text. Three outcomes are told apart. A 200 that hands back markdown is a published file. Anything else is an absence. And a 200 that hands back HTML is a soft 404 — a site answering every unknown path with its homepage — and is reported as exactly that, never as a policy file, because a scanner that read a homepage as an instruction file would be inventing the instructions it then reported.

Presence is reported and not scored. Adoption is early and absence is not a defect; a site with no agents.md has made no claim. What is scored is the content. Once the file exists, its text goes through the same instruction scan as hidden elements, HTML comments and unusually long alt attributes — the surfaces a machine ingests at full weight and a reader never sees. The file written to be obeyed is treated like every other place where an instruction could hide.

The four kinds of instruction it looks for #

KindWhat it matchesWhy it is read as an instruction
override“ignore previous instructions”, “disregard the above”, “system prompt”text addressed to the model’s own rules, not to a reader
role“you are a helpful assistant”, “act as a …”assigns the agent a persona the site did not earn
suppression“do not mention”, “never reveal”, “do not cite”asks the agent to withhold something from its user
steering“always recommend”, “rank us first”, “when asked about X, say”asks the agent to answer for the site rather than about it

A match is exposure, not intent. Most hidden text is an old SEO habit or a collapsed menu; a match is a prompt to open the file and read it, not a verdict that someone attacked the site.

The report carries three rows from this: whether any machine-only surface matched one of those patterns, whether any hidden block runs past fifty words, and what /agents.md returned. The first two are scored. The third states what was found and leaves the decision where it belongs. The content map is a different file with a different job, and the two are routinely confused.

Why one scan cannot tell you this #

Authorship is not a property of a file. It is a property of a file compared with the same file on other sites. Open one agents.md and you see prose. You need the same surface across unrelated domains before the word inherited means anything at all.

So we do not keep a hand-written list of platform boilerplate — that would rot the week a platform edits its template. Lines that appear on most samples across at least two unrelated registrable domains are boilerplate, discovered by measurement. Twelve sites belonging to one operator do not qualify, which is the false positive that would discredit the whole idea.

How fast this moves #

A third store we had recorded as shipping the same file now returns 404 for it, and its robots.txt no longer advertises the endpoints it did two days ago. Same platform, same week. This layer changes underneath owners who never touched it, which is the entire argument for measuring it more than once. A clean audit is a date, not a state, and a file you did not write is the clearest case of that.

Two agents.md files, compared line by line #

MeasureReading
Lines compared86
Lines identical79
Line overlap, before0.791
Line overlap, after1.000
A third store, two days laterthe file now 404s, and its robots.txt no longer advertises the endpoints it did

A file inherited from a template is still a file on your domain, giving instructions in your name.

The file on this domain #

The test cuts both ways. If this scanner flags steering language in other sites’ agents.md, then any steering in our own would be the same defect with our name on it. So ours states what is here — the scan, the aggregate data, the policy, the crawler identity — and asks an agent for nothing. It does not recommend, it does not rank, and it does not point at a third party’s skill file. That is the standard the scanner holds other domains to, and it is easier to hold when the file was written by the people it speaks for.

What to do about it #

Open yourdomain.com/agents.md. If something is there, read it as an instruction aimed at a machine acting for your customer, because that is what it is. Deleting it is rarely right — these files do useful work. Knowing what it says on your behalf is not optional. And if what it says was written by your platform rather than by you, the honest question is not whether the advice is good but whether you would have put your name to it.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Was my agents.md written by me or inherited?

Strip the brand name and domain out and compare it with the same file on unrelated sites. Two files that become identical were both supplied by a platform, not authored.

Why can a single scan not detect an inherited instruction file?

Because inheritance is a property of the file compared with other sites' files. One site in isolation gives you nothing to compare against.

What should I do if my file turns out to be a template?

Read what it actually instructs agents to do on your behalf, then rewrite the parts that do not describe your business.

Where do injected instructions usually live?

HTML comments, hidden or off-screen elements, user-generated fields, and third-party embeds. Anywhere content reaches the page without passing a human.

Is having an agents.md file risky?

Only in the sense that it is public. Anyone can read it, so it should ask for nothing you would not say out loud, and it should not contradict what robots.txt already permits.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works