CrawlCheck

Glossary · area 3 of 19

Machine files

The files addressed to machines rather than readers, and what publishing one does and does not buy.

24 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

24terms in this area
8with a live finding rate

llms.txt

A plain-text file at the root of a site that tells language models what the site is and which pages matter. It is a proposal, not a standard: no engine has committed to reading it, and publishing one is a low-cost bet rather than a ranking action.

Measured: NO_LLMS_TXT 24.4%

llms-full.txt

A proposed companion to llms.txt carrying full content rather than links. Adoption is lower still, and no engine has committed to reading either.

agents.md

A markdown instruction file addressed to autonomous agents rather than crawlers, stating what a site does, what an agent may collect, and what it must never assert. Most files served at this path are platform defaults rather than anything the site owner wrote.

entity map

A pair of files - JSON for machines, an HTML twin for people - listing the entities a site claims to be about. The measurement that matters is not whether the pair exists but whether the two halves agree.

entity map parity

The share of entities named in a site's entity-map JSON that also appear in its rendered HTML twin. Anything present in one and missing from the other is a fact the machine record asserts and no reader can ever see.

machine layer

Everything a site serves to non-human clients: robots.txt, sitemaps, llms.txt, agents.md, entity maps, structured data and response headers. A site can look finished to a visitor and be nearly empty at this layer.

Measured: MACHINE_CHAIN_HEAVY 10.8% · NO_LLMS_TXT 24.4%

well-known URI

A path under /.well-known/ reserved for machine-readable declarations. Agent-facing files are increasingly published there; serving one does not mean anything consumes it.

MCP

Model Context Protocol. A convention for exposing tools and data to assistants. Publishing an MCP endpoint is a capability, not a visibility measure.

agent card

A machine-readable description of what an agent or service can do, served at a well-known path. Early and sparsely adopted.

sitemap index

A sitemap that lists other sitemaps rather than pages. Counting its entries as URLs undercounts a large site by orders of magnitude.

Measured: DECLARED_SITEMAP_BROKEN 1.6% · SITEMAP_UNREADABLE 0.2%

OpenAPI specification

A JSON or YAML description of an API: endpoints, parameters, responses. An agent reads it to find out what it can call and with what arguments. Unlike llms.txt, it is a settled standard with validators.

ai.txt

A proposed file for declaring how AI systems may use a site content, in the same family as llms.txt and Content Signals. Informal proposal, not an enforced standard, and adoption is low enough that absence tells you nothing.

Measured: NO_LLMS_TXT 24.4%

Tool definition

The structured description of a callable function an assistant can invoke: name, parameters, return shape, and a description the model reads to decide when to call it. The description is doing more work than most authors realise.

SKILL.md

A markdown file at the root of a domain, with YAML frontmatter naming the skill, that tells an agent what the site can do for it. A 200 that carries HTML is a soft 404, not a file. Measured in agent surfaces without a score: adoption is early and absence says nothing about the site.

sitemap origin error

A sitemap path answering 500, 502 or 504. A server error is the origin failing, which is neither a refusal nor an absence, so it is reported on its own rather than as a blocked or missing sitemap - one file cannot be both refused and missing. Reported as SITEMAP_ORIGIN_ERROR.

Measured: SITEMAP_ORIGIN_ERROR 0.1%

soft 404 on a machine file

A platform that answers every unknown path with its homepage answers /llms.txt and /sitemap.xml that way too, at HTTP 200. A nonexistent-path control fetched in the same scan identifies the catch-all page, and any machine file that matches it is recorded as absent rather than present or broken.

Measured: SITEMAP_IS_HTML 0.9% · ROBOTS_IS_HTML 0.2%

tdmrep.json

A JSON file at /.well-known/tdmrep.json, defined by the TDM Reservation Protocol from a W3C community group, not a W3C standard. It carries tdm-reservation, a flag stating that text-and-data-mining rights are reserved, and tdm-policy, a link to an ODRL document describing terms. It communicates a legal position under the EU text-and-data-mining exception to operators who read it. By itself it enforces nothing, and this scanner records its presence without treating it as a block, for the same reason it does not treat ai.txt or Content-Signal as one.

TDM exception

Article 4 of the EU copyright directive of 2019: content may be mined for purposes including model training unless the rightsholder has expressly reserved that right in a machine-readable way. It is the legal thread that connects robots.txt directives, Content-Signal, ai.txt and tdmrep.json to something with force: a reservation expressed in one of them can matter in a European court. Whether a given operator reads any of them before fetching remains a policy question, and an expressed reservation is evidence of intent, not a guarantee of exclusion.

security.txt

A file at /.well-known/security.txt, defined by RFC 9116, giving a contact for reporting vulnerabilities and an expiry date for that information. It is included here as the contrast case: a well-known machine file with an actual standard behind it, which llms.txt, ai.txt and agents.md do not have. It supports exactly one claim, that a reporting route is published; its presence says nothing about the site's security, and an expired one is worse than none because it is a promise nobody kept.

Content Signals

A robots.txt extension proposed in 2025 in which a site states, per path, whether its content may be used for search, for AI input at answer time, and for AI training, as three separate signals. It is a declaration crawlers may read; nothing enforces it, and a crawler that ignores robots.txt ignores this too.

HTTP Message Signatures

A standard, RFC 9421, for signing HTTP requests so a server can verify who sent them. Web Bot Auth applies it to crawlers: the operator publishes a key directory, signs each request, and a site can verify the signature instead of trusting a user-agent string or an IP list. It answers who sent the request, not whether they may have what they asked for.

changefreq and priority

Two optional sitemap fields, one guessing how often a page changes and one ranking pages against each other. Google has said it ignores both; other engines mostly do. They cost nothing to include and buy nothing when included, and a priority of 1.0 on every URL is the same as none.

staging leak

A production page whose source references the site's staging hostname on a managed host, usually because a search-and-replace missed a serialised option, page-builder JSON or a theme file. A crawler that reads the page learns the host, which is typically a full, indexable and out-of-date copy of the site.

Measured: STAGING_LEAK 0.1%

host duplicate

The www and bare forms of a domain both answering 200 with the same page, neither redirecting to the other. Two addresses for one page are two candidates for one query, and the canonical tag, if present, is a hint the engine may not take. The fix is one host that answers and every other host that redirects to it.

Measured: HOST_DUPLICATE_200 3.3%

← Crawlers and access  ·  Entities and schema →

All 668 terms across 19 areas.