CrawlCheck

Reference

AI and LLM ranking terms, defined with their limits

In one sentence

The CrawlCheck glossary defines 179 AI-visibility, GEO, AEO and AI-crawler terms across eight areas — from AI visibility score, llms.txt and entity graph to verified crawler, RAG and grounding — and every definition states what the term can support and what it cannot, because most confusion in this field is a word doing more work than the evidence behind it.

179 terms across 8 areas. Every definition states what the term can support and what it cannot, because most of the confusion in this field is a word doing more work than the evidence behind it.

Written against what this scanner measures. Where a term is commonly used to mean more than it can carry — a citation, a forged crawler, a visibility score, a passing file — the definition says so rather than repeating the looser sense. Several words here mean two different things depending on who is speaking; those entries name the collision instead of picking a side.

How to read these

Three habits carry most of the value. Ask what the denominator is — a rate without one is decoration. Ask whether absence was measured or assumed — unverifiable is not the same as forged, and not found is not the same as not there. Ask what the claim would look like if it were wrong — a term that cannot fail a test is not measuring anything.

The eight areas, grouped by the stage they belong toThree stages in sequence. Reach covers crawlers and access, and delivery and rendering. Read covers machine files, entities and schema, and indexing and discovery. Quote covers answer engines and retrieval, measurement and evidence, and proof and provenance. A term is only useful once the stage before it holds.1 · Reach2 · Read3 · Quotecan a client get the bytescan it parse what it gotcan an engine safely repeat itCrawlers and accessDelivery and renderingMachine filesEntities and schemaIndexing and discoveryAnswer engines, retrievalMeasurement and evidenceProof and provenanceA term in stage 3 is worth nothing if stage 1 never happened: a page that was never fetchedcannot be misread. Most arguments about AI visibility are stage-3 arguments about siteswith a stage-1 problem. Each definition below sits in the section that names its stage.

Every term, A to Z

202 interstitial · @context · accessibility structure · accessible name · AEO · agent card · agentic knowledge graph · agentic traffic · agents.md · AI Mode · AI opt-out · AI Overview · AI referral attribution · AI visibility score · ai.txt · alt attribute · anchoring · answer engine · ASN blocking · blocked render resource · BM25 · bot management platform · brand mention · Bytespider · C2PA · canonical · case consistency · CAWG training and data mining assertion · CCBot · challenge page at 200 · change receipt · Chunk boundary · chunking · citation · ClaudeBot vs Claude-User · cloaking · cohort asymmetry · Conditional request · Content hash · content-addressed storage · Content-Signal · context window · Core Web Vitals · cosmetic row · crawl budget · crawler classification · Crawler trap · cross-encoder · CrUX · dangling @id · dated series · DefinedTerm · denominator note · dense retrieval · DOM depth · double-counted defect · duplicate id · Dynamic rendering · E-E-A-T · E-E-A-T proxies · edge cache pinning · Edge transformation · embedding · entity collision · entity disambiguation · entity graph · entity map · entity map parity · entity resolution · evidence level · faceted navigation · false positive · FAQPage · FCrDNS · fetcher disagreement · field data · finding age · forged crawler identity · form control name · founder edge · foundingDate · gated row · GEO · Google-Agent · Google-Extended · GPTBot vs OAI-SearchBot · graded through a wall · grounding · hallucinated attribution · hallucination · homepage refused · Hybrid search · Hydration mismatch · index coverage · IndexNow · information gain · JSON-LD · knowledge graph · knowsAbout · landmark region · llms-full.txt · llms.txt · local citation · LocalBusiness · log file analysis · machine layer · main landmark · Manifest tampering · MCP · microformats (h-card) · named entity recognition · NAP · noindex · nonexistent-path control · off-host redirect · OpenAPI specification · OpenTimestamps · operator feed · Organization schema · orphan page · own-host loop · Passage ranking · payload ratio · percentile rank · PerplexityBot vs Perplexity-User · Person schema · placeholder in record · population baseline · prompt injection · prompt set · Query fan-out · RAG · rate limiting · RDF · RDFa and Microdata · Re-ranking · reciprocated sameAs · refusal · regression · render dependence · render dependence gap · residential proxy · retrieval · retrieval crawler · robots.txt · row review · sameAs · score version · security.txt · self-audit · self-audit invariant · share of voice · sitemap index · sitemap origin error · SKILL.md · skip link · Soft 404 · soft 404 on a machine file · stable @id · structured data · structured data validation · TDM exception · tdmrep.json · token · Tool definition · topical authority · training data · triple · TTFB · uniform refusal · unreciprocated profile · unscored measurement · unverifiable · user-agent group · user-triggered fetch · verified crawler · well-known URI · Wikidata · X-Robots-Tag

Answer engines and retrieval

How a generated answer gets made, and what your page is competing for inside it.

AEO

Optimising to be quoted inside a generated answer rather than ranked in a list of links. The practices overlap heavily with SEO; what differs is the outcome measured. No engine publishes a ranking function for answers, so any claim to optimise for one is inference from correlation.

GEO

A newer name for the same activity as AEO. The two terms are used interchangeably in practice and no consistent technical distinction has been established between them.

AI visibility score

A composite number sold by several tools to represent how visible a brand is to language models. There is no shared definition, no published methodology across vendors, and no engine endorses any of them. Two tools can score the same site 40 and 85 and both be internally consistent.

share of voice

The proportion of sampled answers mentioning your brand versus named competitors. Meaningful only alongside its sample size, its prompt set and its date. A share of voice quoted without those three is a number without a denominator.

prompt set

The fixed list of questions a monitoring tool sends to models to see who gets mentioned. Results move when the prompt set changes, so a prompt set that is not published cannot be audited or reproduced.

brand mention

A model naming your brand in an answer. Distinct from a citation: a mention carries no link and no source attribution, so it cannot be traced back to a page you control.

RAG

Retrieval-augmented generation. The model retrieves documents at question time and writes an answer grounded in them. Under RAG your page competes to be RETRIEVED, which is a different problem from being in the training data.

grounding

Constraining a generated answer to retrieved source material. Grounding reduces invention but does not eliminate it: a grounded answer can still misattribute which source said what.

training data

Text absorbed during model training. Content in training data may influence answers without ever being cited, and cannot be updated or withdrawn once a model is trained. This is why blocking training crawlers and allowing retrieval crawlers is a coherent policy rather than a contradiction.

embedding

A numeric vector representing text, used to find passages similar in meaning rather than in wording. Nothing about a page is retrievable by embedding if the page was never fetched in the first place.

chunking

Splitting a page into passages before indexing. A page can be retrieved as one passage that reads badly out of context, which is why a self-contained paragraph is easier to quote than one that depends on the paragraph above it.

context window

The amount of text a model can consider at once. A page that exceeds what a retriever will pass along is truncated silently, and the part that survives is not necessarily the part you would have chosen.

token

The unit a model reads, roughly a word fragment. Token counts matter because markup, navigation and boilerplate consume the same budget as the content they surround.

hallucination

A confident statement a model generates that is not supported by any source. Distinct from hallucinated attribution, where the claim is right but the credited source is wrong.

refusal

An engine declining to answer. A refusal is not evidence about your site: it is evidence about the engine policy at that moment, and treating it as a site defect inverts cause and effect.

AI Overview

Google generated answers shown above the results. Appearance is not controllable and not guaranteed, and the sources shown change between identical searches, so absence on one sample proves nothing.

answer engine

A system that returns a composed answer with citations rather than a list of links. Its retrieval step is what a site can influence; its generation step is not.

citation

A source an answer engine names when composing an answer. A citation proves a document was retrieved and used once; it does not describe the model's ranking and it decays as weighting changes.

retrieval

The step in which an answer engine selects documents to compose from. Everything a site controls acts here, which is why machine-readable clarity outperforms persuasion.

hallucinated attribution

An answer that credits a claim to a source which does not contain it. The defence is not more content but content that states its own scope, dates and limits.

Running lexical search (BM25 or similar) and dense vector search together at the retrieval step, then merging the results. It exists because pure semantic search misses exact strings — part numbers, model names, a company name that looks like an ordinary word.

Re-ranking

A second scoring pass, usually a cross-encoder, that reorders the chunks retrieval returned before the generator sees them. A page can win retrieval and still be dropped here, which is why "we were retrieved" and "we were cited" are different claims.

Passage ranking

Scoring individual text segments rather than whole pages. It is why one clearly written paragraph can be cited from a page that is otherwise thin, and why page-level authority metrics predict citation poorly.

Query fan-out

An engine rewriting one user prompt into several internal queries before searching. Optimising for the prompt a person typed ignores the queries the system actually ran, and neither we nor any vendor can observe those directly.

Chunk boundary

The point where text is split to fit a token limit. A split between a premise and its conclusion leaves both halves less usable, and the split is made by the engine, not by you — which is why short self-contained passages survive it better.

BM25

The lexical half of hybrid search: a ranking function that scores a document by how often a query term appears in it, discounted by how common the term is across the corpus and by document length. It finds exact strings that embedding similarity blurs, a part number, a company name that is also an ordinary word. It cannot find a paraphrase, which is why retrieval systems pair it with dense retrieval rather than choosing one.

cross-encoder

The model class usually behind re-ranking. It reads the query and a candidate passage together, in one pass, and emits a relevance score, which is more accurate than comparing two separately computed vectors and far slower. It runs only on the short list retrieval has already produced. A passage that retrieval never surfaced is never scored by it, so improving how a page reads helps here only after the page is findable at all.

AI Mode

Google's conversational search surface, a separate tab and mode from the AI Overview that appears above ordinary results. Both draw on the Search index, but they fan a query out differently, cite differently, and refresh on different cycles, so a page can be cited in one and absent from the other. We appear in Google's AI results is therefore two claims, and a measurement that does not say which surface it read is not reproducible.

information gain

A ranking concept with a Google patent lineage: a document is scored for what it adds beyond the documents the user, or the retrieval set, has already seen. It explains a pattern visible from outside, that a page restating the consensus answer loses to a page contributing a figure, a method, or a dated observation nobody else has. The score itself is not observable from outside an engine, so it is a working explanation, not a measurement, and this scanner does not claim to compute it.

dense retrieval

The semantic half of hybrid search: a query and each candidate passage are turned into vectors by an embedding model and matched by similarity, so a paraphrase can be found. Embedding names the vector; dense retrieval names the search step those vectors exist for. It misses exact strings that BM25 finds, it inherits every bias of the model that made the vectors, and a passage that was chunked badly is a bad vector however well it was written.

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

retrieval crawler

A crawler that fetches pages to answer a question being asked right now. Distinct from a training crawler by purpose, not by capability, and the two are often operated by the same company under different user-agents.

user-triggered fetch

A request made because a person asked the assistant about a specific URL. Volume reflects user curiosity rather than crawl policy, so it is a poor proxy for how well a site is indexed.

verified crawler

A request whose claimed crawler identity is confirmed by checking the source address against the operator's own published IP ranges. Reverse DNS and user-agent matching are weaker: only the operator's feed is authoritative. measured on /ai-crawlers

forged crawler identity

A request whose user-agent names a known crawler while its source address falls outside the ranges that operator publishes. Verified against a published feed it is a fact, not an accusation; without a feed it is unverifiable. measured on /tools

unverifiable

The third outcome, and the one most tools omit. A profile behind a login wall, a 429, or a shell page under 512 bytes can be neither confirmed nor accused, and folding it into either bucket produces a number that overstates or understates by construction.

robots.txt

A file declaring which paths each user-agent may fetch. It controls access, not indexing, and a 403 on robots.txt itself is read by most compliant crawlers as disallow-everything.

user-agent group

A block of robots.txt rules addressed to named agents. Most parsers apply only the most specific matching group, so naming an agent to allow it can silently remove every rule it would otherwise have followed.

AI opt-out

Any published instruction asking that a site's content not be used for model training. It is a request rather than a control, and whether it is honoured varies by operator.

Content-Signal

A robots.txt directive expressing separate permissions for search indexing, AI input and AI training. It communicates intent to operators who choose to honour it and carries no enforcement of its own.

crawl budget

The effort a crawler will spend on a site before moving on. Wasting it on redirect chains and unreachable URLs is measurable; the budget itself is not published by any operator.

cloaking

Serving materially different content to a crawler than to a browser. A server refusing a datacentre address while serving a residential one is not cloaking - it is IP-based access control, and calling it cloaking is a false accusation.

challenge page at 200

A bot-protection interstitial returned with a success status. Any validator that trusts the status code records the file as present and readable when nothing readable was served.

uniform refusal

An origin answering the homepage and a path that cannot exist the same way: same status, same size, or the same challenge or block page. A site does not serve a missing page identically to its home page, so the answers describe a wall rather than the site. CrawlCheck records exactly one finding for such a scan and grades nothing; a score produced through a wall would be a number about someone else's product.

nonexistent-path control

A request for a randomised URL that no site would ever publish, made alongside the real requests. It is the control that tells a scanner whether it reached the site at all: a real site answers it differently from its homepage. Tightened on 1 September 2026 to catch walls that answer with non-HTML, a challenge page of any size, or a block page byte-identical across paths.

blocked render resource

A stylesheet or script the homepage references that the site's own robots.txt disallows to every crawler. A rendering crawler obeys the rule and draws the page without it, so layout and any script-inserted content are judged from a page the site never intended to show. Reported as RENDER_RESOURCE_BLOCKED from the * group only.

off-host redirect

A homepage that answers with a redirect to a hostname outside its own www/apex pair, such as a regional storefront or a checkout host. The scanner does not follow it: whatever a crawler reads on the other host is attributed to that host, and the requested domain reads as a site with no homepage. Reported as HOMEPAGE_REDIRECTS_OFF_HOST.

homepage refused

A homepage answering 401, 403, 429 or similar to this client while robots.txt is served normally. The machine layer stays scored; every section that needs the homepage is unmeasured rather than scored, and the report says why. It is marked unconfirmed until checked from a second address, because a datacentre refusal may be about the scanner rather than about crawlers. Reported as HOMEPAGE_REFUSED.

FCrDNS

Forward-confirmed reverse DNS: resolve the requesting IP to a hostname, then resolve that hostname back and check it returns the same IP. It is the verification method for crawlers that publish no IP list, and it is the only honest way to test a Googlebot claim.

Crawler trap

An unbounded set of generated URLs — faceted filters, infinite calendars, recursive relative paths — that a crawler can follow forever. It consumes crawl allocation without exposing new content.

Conditional request

A fetch carrying If-None-Match or If-Modified-Since, which lets the server answer 304 Not Modified instead of resending the body. A site that never answers 304 pays full bandwidth for every re-crawl.

ASN blocking

Refusing traffic by hosting network rather than by identity. It blocks retrieval agents running on cloud infrastructure — which is most of them — while leaving residential proxies untouched, so it usually costs more visibility than it prevents scraping.

operator feed

The publisher-side list of source addresses a crawler operator maintains, usually an IP-range JSON, a reverse-DNS suffix, or both. It is the only authoritative basis for calling a hit verified or forged: a request from an address in the feed is the operator's, a request outside it is not, whatever the user-agent string says. Meta, Apple and Amazon publish only a documentation page, so their hits can be checked by reverse DNS but not by range. A stale feed indicts the operator's publishing, not the site being measured.

residential proxy

A request relayed through a consumer ISP address so that it is indistinguishable, at the network layer, from a person at home. It defeats ASN blocking and it defeats feed verification, because the address is in nobody's operator feed and belongs to nobody's data centre. Any forged-crawler rate computed against operator feeds is therefore a floor: the forgeries that used residential addresses were never counted, and no server-side method can count them.

Google-Extended

The robots.txt token that controls whether Google may use a site's content to train and ground its Gemini models. It is not a crawler; Googlebot still fetches the page. Blocking it does not remove a site from Google Search, from Discover, or from the retrieval path behind AI Overviews and AI Mode, which run on the Search index. It is the most commonly misread token on the web: sites block it expecting to leave AI answers and instead leave only the training set.

Google-Agent

A user-triggered fetcher Google added to its documentation on 20 March 2026 for agents that run on Google infrastructure and browse on a person's behalf, Project Mariner among them. It carries a Chrome-like user-agent string with the token inside it and publishes its own range file, user-triggered-agents.json, separate from the other fetcher feeds. Google classes it as a user proxy, so it generally ignores robots.txt. It is the concrete case where the caveat on user-triggered fetch stops being theoretical: a rule addressed to it is a request the fetcher is documented not to read.

GPTBot vs OAI-SearchBot

Two OpenAI crawlers with one operator and different purposes. GPTBot collects for model training; OAI-SearchBot fetches for live search retrieval and is what a ChatGPT search citation depends on; a third identity, ChatGPT-User, fetches when a person asks about a URL. Each has its own range in OpenAI's published feed and its own robots.txt token. Blocking one says nothing about the others, which is why allow and block decisions are per agent, not per company, and why a rule addressed to GPTBot alone leaves search retrieval untouched.

ClaudeBot vs Claude-User

The same split for Anthropic. ClaudeBot crawls autonomously to improve models; Claude-User fetches a page because a person asked Claude about it, and Claude-SearchBot fetches to build and refresh search results. All three verify against Anthropic's published IP ranges. Blocking ClaudeBot does not stop a user-triggered fetch, and counting a Claude-User hit as a crawl inflates a crawl figure with visits that were, functionally, referrals.

PerplexityBot vs Perplexity-User

The same split for Perplexity. PerplexityBot indexes for the answer engine and honours robots.txt; Perplexity-User fetches a page when a person's query requires it and, by Perplexity's own documentation, generally does not honour robots.txt because the request is treated as the user's. Both publish IP ranges. A site that sees Perplexity fetches after blocking PerplexityBot is usually seeing the second identity, not a violation by the first.

CCBot

The Common Crawl crawler. Its open corpus is downloaded and used by many model builders, so a page can enter training sets without ever being fetched by any model operator's own crawler. That is the case that breaks the sentence we blocked GPTBot, so we are out of the training data: the copy that trained the model may have been taken by a nonprofit archive years earlier. CCBot honours robots.txt and publishes its identity, so exclusion is possible; it is only retroactive removal that is not.

Bytespider

ByteDance's training crawler, reported over several years by publishers and by edge providers to fetch without regard to robots.txt exclusions. It is the standing example that a robots.txt rule is a request, honoured by operators who choose to honour it and enforced by nothing. The practical consequence for measurement is ordering: verify who actually fetched before attributing volume to a name, because a crawler that ignores rules is also the easiest name to forge.

agentic traffic

Requests from agentic browsers, Comet, Atlas and their successors, in which a model drives a real browser session on a person's behalf. They arrive with an ordinary browser user-agent, cookies, and human-looking session patterns, from consumer addresses. No user-agent match, operator feed or reverse-DNS check can classify them, so they are invisible to every verification method in this glossary. This is a definable limit rather than a solved problem: a verified count is a count of what identified itself, and this traffic does not.

crawler classification

The model edge providers adopted in 2026, Cloudflare from July with new-domain defaults changing on 15 September: verification confirms only who a crawler is, and access is then decided per behaviour class, Search, Agent, or Training, rather than per operator. A multi-purpose crawler is judged under every class it belongs to, which is how a rule that blocks Training can block a search crawler that also trains. It is the industry name for the identity-versus-permission distinction that verified crawler already implies. The classes are the provider's judgment of a crawler's purpose, and that judgment is not published in a form a site can audit.

Machine files

The files addressed to machines rather than readers, and what publishing one does and does not buy.

llms.txt

A plain-text file at the root of a site that tells language models what the site is and which pages matter. It is a proposal, not a standard: no engine has committed to reading it, and publishing one is a low-cost bet rather than a ranking action. measured on /data

llms-full.txt

A proposed companion to llms.txt carrying full content rather than links. Adoption is lower still, and no engine has committed to reading either.

agents.md

A markdown instruction file addressed to autonomous agents rather than crawlers, stating what a site does, what an agent may collect, and what it must never assert. Most files served at this path are platform defaults rather than anything the site owner wrote.

entity map

A pair of files - JSON for machines, an HTML twin for people - listing the entities a site claims to be about. The measurement that matters is not whether the pair exists but whether the two halves agree. measured on /data

entity map parity

The share of entities named in a site's entity-map JSON that also appear in its rendered HTML twin. Anything present in one and missing from the other is a fact the machine record asserts and no reader can ever see.

machine layer

Everything a site serves to non-human clients: robots.txt, sitemaps, llms.txt, agents.md, entity maps, structured data and response headers. A site can look finished to a visitor and be nearly empty at this layer.

well-known URI

A path under /.well-known/ reserved for machine-readable declarations. Agent-facing files are increasingly published there; serving one does not mean anything consumes it.

MCP

Model Context Protocol. A convention for exposing tools and data to assistants. Publishing an MCP endpoint is a capability, not a visibility measure.

agent card

A machine-readable description of what an agent or service can do, served at a well-known path. Early and sparsely adopted.

sitemap index

A sitemap that lists other sitemaps rather than pages. Counting its entries as URLs undercounts a large site by orders of magnitude.

OpenAPI specification

A JSON or YAML description of an API: endpoints, parameters, responses. An agent reads it to find out what it can call and with what arguments. Unlike llms.txt, it is a settled standard with validators.

ai.txt

A proposed file for declaring how AI systems may use a site content, in the same family as llms.txt and Content Signals. Informal proposal, not an enforced standard, and adoption is low enough that absence tells you nothing.

Tool definition

The structured description of a callable function an assistant can invoke: name, parameters, return shape, and a description the model reads to decide when to call it. The description is doing more work than most authors realise.

SKILL.md

A markdown file at the root of a domain, with YAML frontmatter naming the skill, that tells an agent what the site can do for it. A 200 that carries HTML is a soft 404, not a file. Measured in agent surfaces without a score: adoption is early and absence says nothing about the site.

sitemap origin error

A sitemap path answering 500, 502 or 504. A server error is the origin failing, which is neither a refusal nor an absence, so it is reported on its own rather than as a blocked or missing sitemap - one file cannot be both refused and missing. Reported as SITEMAP_ORIGIN_ERROR.

soft 404 on a machine file

A platform that answers every unknown path with its homepage answers /llms.txt and /sitemap.xml that way too, at HTTP 200. A nonexistent-path control fetched in the same scan identifies the catch-all page, and any machine file that matches it is recorded as absent rather than present or broken.

tdmrep.json

A JSON file at /.well-known/tdmrep.json, defined by the TDM Reservation Protocol from a W3C community group, not a W3C standard. It carries tdm-reservation, a flag stating that text-and-data-mining rights are reserved, and tdm-policy, a link to an ODRL document describing terms. It communicates a legal position under the EU text-and-data-mining exception to operators who read it. By itself it enforces nothing, and this scanner records its presence without treating it as a block, for the same reason it does not treat ai.txt or Content-Signal as one.

TDM exception

Article 4 of the EU copyright directive of 2019: content may be mined for purposes including model training unless the rightsholder has expressly reserved that right in a machine-readable way. It is the legal thread that connects robots.txt directives, Content-Signal, ai.txt and tdmrep.json to something with force: a reservation expressed in one of them can matter in a European court. Whether a given operator reads any of them before fetching remains a policy question, and an expressed reservation is evidence of intent, not a guarantee of exclusion.

security.txt

A file at /.well-known/security.txt, defined by RFC 9116, giving a contact for reporting vulnerabilities and an expiry date for that information. It is included here as the contrast case: a well-known machine file with an actual standard behind it, which llms.txt, ai.txt and agents.md do not have. It supports exactly one claim, that a reporting route is published; its presence says nothing about the site's security, and an expired one is worse than none because it is a promise nobody kept.

Entities and schema

Declaring who you are in a form a resolver can use without guessing.

structured data

Machine-readable statements about a page, usually JSON-LD, describing what the page is and what it is about. Validity is cheap to check; agreement with the visible page is the part that is usually wrong.

JSON-LD

The JSON serialisation of linked data, embedded in a script tag. It is the format search and answer engines read most reliably, which also means an error in it is an error stated confidently to machines.

stable @id

A durable identifier that lets separate statements be understood as being about the same thing. Two records sharing an @id are the same record twice; without one, every page's claims start again from nothing.

entity collision

Two distinct entities in a site's structured data sharing an identifier, or one entity split across identifiers. Either way a machine reading the graph cannot tell how many things are being described.

reciprocated sameAs

A sameAs link that the destination profile returns. Counting sameAs links measures how many claims a site makes about its own identity; fetching each destination and looking for the domain measures how many of those claims anyone else confirms.

sameAs

A schema property linking an entity to its profiles elsewhere. Its value is entirely in whether the target links back; without that it is a claim about yourself.

knowledge graph

A store of entities and relationships an engine keeps independently of any page. Publishing correct markup makes a claim available for reconciliation; it does not put anything into a graph by itself.

entity resolution

Deciding whether two records describe the same thing. Everything a site can do to help is about giving one identity one stable identifier and never contradicting itself.

entity disambiguation

Distinguishing your entity from others sharing a name. A shared name with no distinguishing identifier is the ordinary case for local businesses, not an edge case.

Wikidata

An open structured database used as an authority link target. A sameAs pointing at Wikidata is only corroboration if the Wikidata item points back; an unreciprocated link is a claim, not a confirmation.

NAP

Name, address and phone number. Its value is consistency across every place the business appears, so the failure mode worth measuring is disagreement between sources rather than absence from any one.

local citation

A directory or profile listing a business. Note the collision: in the answer-engine sense a citation is a source credited in an answer. This glossary keeps them separate because conflating them makes both numbers meaningless.

Organization schema

Markup declaring the publishing entity. The single most reusable node on a site, because everything else can reference it by identifier rather than repeating the same facts.

LocalBusiness

Markup for a business with a service area or premises. Coordinates and a street address that disagree are a defect a resolver cannot arbitrate.

FAQPage

Schema marking question-and-answer content. The answers must match what a visitor sees; declaring answers not present on the page is the ordinary way this markup becomes a liability.

structured data validation

Checking markup parses and matches the visible page. Valid markup that contradicts the visible text is a worse position than no markup, because it makes a confident wrong claim available for reuse.

@context

The JSON-LD declaration binding your keys to an external vocabulary. Without a valid context, name and description are just strings — the linked-data meaning is carried by the context, not by the key.

agentic knowledge graph

A knowledge graph that agents read from and write to, rather than one a person curates and queries. Two senses are in circulation and they are not the same claim: the practical one, where a graph stores entities and relations that a retrieval agent traverses instead of, or alongside, embedding search; and a speculative one, where the graph is generated by a model as it reasons and passed between agents as a reasoning artifact. The first is a database design and is in production; the second is a research and blog-post idea with no standard behind it. Neither is something a website publishes. Your site can contribute nodes to somebody else graph by declaring stable identifiers and reciprocated sameAs edges, but you do not control what the graph makes of them.

microformats (h-card)

Contact and identity markup carried in HTML class attributes rather than in JSON-LD: a name, organisation, email, URL and locality that any parser can read from the page itself. Measured and shown in the schema section, and compared against the NAP where both exist. Not scored: adoption outside a few communities is a rounding error, and absence is not a defect.

RDF

The Resource Description Framework: facts expressed as subject, predicate, object statements (triples) with resolvable identifiers. JSON-LD is one serialisation of it, so every schema.org block on a page is already a small RDF graph, and it can be checked as one - whether its identifiers resolve, whether references point at nodes that exist, whether it parses at all.

triple

One statement in an RDF graph: this subject has this property with this value. A JSON-LD node with a name, a URL and a type is three triples. Counting them is the honest size of what a linked-data reader receives; a page can carry a large script block and declare almost nothing.

dangling @id

A reference to an @id that no node on the page declares. The reference is syntactically valid, so validators pass it, and it is a statement about nothing: the resolver drops the edge and the entity loses the relation it was meant to carry.

RDFa and Microdata

Two ways of embedding structured data in HTML attributes rather than a script block. Read by some parsers, ignored by others, and a second place for the facts to disagree with the JSON-LD. Counted alongside JSON-LD and not scored there, because one serialisation kept accurate is worth more than three that drift. When a page declares no JSON-LD node at all, its schema.org Microdata and RDFa are read as the page's graph and every entity check runs on them (from score version 20).

founder edge

The founder or employee property on an Organization node, pointing at a Person by @id. It is the statement that connects the business graph to the people graph. The target must be declared on the same page or the edge dangles; a reference is not a declaration.

foundingDate

The date property on an Organization node. It is the one claim behind every years-of-experience counter that a reader can check against a registry or a domain record, so it should be a date the operator can defend: incorporation, or the day the site first went live.

knowsAbout

A property on a Person or Organization listing the subjects the entity claims expertise in, ideally as DefinedTerm nodes with a sameAs to a public identifier. Limited to what the site's own content evidences, it is one of the few machine-readable expertise statements; padded, it is a list of keywords nobody will believe.

unreciprocated profile

A sameAs target that is reachable and does not link back. It says nothing about whether the profile is real; some platforms never link a website out. It is shown, not failed. Only a dead target, one that does not answer from a client that can read it, fails a corroboration row.

entity graph

The web of nodes and typed edges a resolver builds from your declarations plus everyone else's: your Organization node, the Person it names as founder, the Wikidata item that names you back. It is distinct from the entity map, which is the file pair a site publishes, and from a knowledge graph, which is the store an engine keeps. A site contributes nodes and edges to a graph; it does not own the graph, and a claim nobody else corroborates is an edge with one end.

DefinedTerm

The schema.org type for a term that has its own identity, used inside knowsAbout so that a claimed subject carries a sameAs to a public identifier instead of sitting there as a bare string. It gives a resolver something to match against Wikidata or a vocabulary. It is a claim about a subject, not about the organisation: a DefinedTerm pointing at a real Wikidata item does not corroborate the organisation's own identity, and a scanner that counts it as an unreciprocated sameAs on the business is reading the wrong node.

Person schema

The node that Organization, LocalBusiness, founder edge and author references all point at, and the authorship half of an entity graph. One Person, one stable @id declared on every page that references it, a name, a jobTitle, a url, and a sameAs to at least one profile that links back. It can support a claim that a named human stands behind the site. It cannot support expertise: a Person node with a credentials list is a declaration, and the scanner's E-E-A-T proxy rows read only whether it is declared and corroborated, not whether it is true.

named entity recognition

How a model extracts people, organisations, places and products from unmarked prose when no structured data exists. It is why a page with no schema is still partly understood, and why structured data is an advantage rather than a requirement. What it cannot give is a stable identity: a name found in text is a string, matched by guess to whichever entity seems likeliest, and two organisations with similar names are confused exactly where a stable @id would have separated them.

Indexing and discovery

Whether a page can be found and kept, separately from whether it can be fetched.

index coverage

How many of a site pages an engine has actually indexed. Distinct from sitemap coverage, which is only what the site declares. The two are compared, never substituted.

canonical

A declaration of which URL is the preferred version. A canonical pointing away from a page is a request not to index that page, and a sitemap declaring a URL whose canonical points elsewhere is a contradiction the site is publishing about itself.

noindex

An instruction to keep a page out of search results, sent as a header or a meta tag. A crawler obeys either, so setting it in one place and not the other still removes the page.

X-Robots-Tag

An HTTP header carrying indexing directives. Invisible in the page source, which is why a page can look perfectly indexable in a browser while being excluded.

orphan page

A page no internal link points to. Findable by a crawler only if it is declared in a sitemap, and findable by nobody if it is in neither.

E-E-A-T

Experience, expertise, authoritativeness, trustworthiness. A framework from Google rater guidelines, not a measurable score. Nothing on a page can be tested to return an E-E-A-T value.

E-E-A-T proxies

The machine-readable statements a site makes that a reader could use as evidence of experience, expertise, authoritativeness and trust: a named person and what qualifies them, corroborating profiles, typed authors and dates on articles, who runs the organisation, how to reach it, and whether the about, contact and policy pages are linked. Each is measurable; none of them IS the quality it stands in for.

topical authority

The idea that covering a subject thoroughly improves standing across it. Widely held and hard to test on one site, because coverage and quality usually change together.

Soft 404

A page that answers HTTP 200 while telling a human the content is missing. The status line and the content disagree, and the crawler believes the status line until something else corrects it.

case consistency

Whether the site's own links and sitemap entries agree on how a URL is written: lowercase paths, one trailing-slash form, one host form, no navigation through query strings. To a crawler /About and /about are two URLs, and a site that links to both has published every page twice. Scored from the hrefs and sitemap entries the site itself serves, never from the paths the scanner chose to request.

accessibility structure

What can be read about accessibility from the document a server returns, without rendering it: a lang attribute, a title, alt on images, names on form controls and buttons, titles on frames, a main landmark, a skip link or landmarks, unique ids. A screen reader depends on all of these before contrast or focus order ever matters. Measured here; a WCAG audit is not.

landmark region

An HTML region a screen-reader user can jump to directly: main, nav, header, footer, aside, or an element with the matching role. One main landmark is the target for skipping the navigation; two is as unusable as none.

accessible name

The text assistive technology announces for a control: visible text, a label, aria-label or aria-labelledby, an image's alt, a frame's title. A button with none is announced as "button" and nothing else; an input with none as "edit text". A placeholder is not a name.

A link, usually first in the document, that jumps past the navigation to the main content. Without it, or without landmarks that serve the same purpose, every page starts by tabbing through the whole menu.

main landmark

The element, main or role=main, that marks the primary content of a page so a screen reader can jump to it. Exactly one per page. A theme that prints it on its own templates often prints nothing on a page builder's, so the pages that matter are the ones to check.

alt attribute

The text alternative on an image. Present and empty marks the image as decorative and is correct; present and descriptive is read aloud; absent entirely makes a screen reader announce the file name. The defect is absence, and a check that fails on empty alt reports every decorative icon as a fault.

form control name

What assistive technology announces for an input: a label element pointing at its id, a wrapping label, aria-label, aria-labelledby, or a title. A placeholder is not a name; it disappears on the first keystroke and is not announced as the field's purpose.

duplicate id

Two elements on one page with the same id attribute. Labels, aria-labelledby, aria-describedby and skip links all resolve by id and land on the first match, which is usually the wrong element. Invalid HTML that validators tolerate and screen readers do not.

IndexNow

A push protocol in which the site tells participating search engines that a URL changed, instead of waiting to be recrawled: a key file on the host, then one request per changed URL. It shortens the delay for engines that participate, chiefly Bing, Yandex, Naver and Seznam. Google does not participate, and the AI answer engines largely do not, so a submission proves the site announced a change; it does not put the page in front of any answer engine.

faceted navigation

Filter combinations, size, colour, price, sort order, that each generate a URL, so that a catalogue of a hundred products can expose millions of addresses. It is the ordinary-case cause of a crawler trap, and the usual reason crawl budget is spent on pages that all show the same items. The fix is a rule about which combinations are canonical, which are noindexed, and which are blocked; case consistency in the parameter names is the related failure, since ?Color= and ?color= are two more URLs.

Delivery and rendering

What arrives on the wire, and what only exists after something else runs.

payload ratio

The share of a delivered page that is text a machine can quote, against markup, scripts and styling that it cannot. A page can be large and still carry almost nothing an answer engine can use.

render dependence

How much of a page's meaning exists only after JavaScript runs. Crawlers that do not execute scripts see whatever the server sent, which on some sites is a shell.

render dependence gap

The difference between what is in the HTML and what appears after scripts run. A non-rendering crawler receives the first, which is why a service list that only exists after JavaScript can be invisible to a fetcher that reached the page successfully.

edge cache pinning

A cached response continuing to be served after the origin has changed. A purge API returning 200 is a statement about the API call, not evidence that anything was evicted.

TTFB

Time to first byte: how long a server takes to begin answering. It is the speed measurement that matters most to a crawler, because a crawler that times out records nothing at all.

Core Web Vitals

Google's field measurements of loading, interactivity and layout stability, drawn from real Chrome users. Two of the three describe rendering and interaction, which no AI crawler performs.

CrUX

The Chrome User Experience Report, which publishes real-user performance for origins with enough Chrome traffic. Absence of data is a statement about traffic volume, never about speed.

field data

Performance measured from real visits, as opposed to a laboratory run. It is the more honest number and it does not exist for most small sites.

prompt injection

Text on a page written to instruct a model reading it rather than the person. Relevant to publishers because instructions can arrive on a domain the owner did not write, through user content or a compromised dependency.

Hydration mismatch

A difference between the server-rendered HTML and the DOM the client script produces. A fetcher reading only the first response records the server version, which may not be what any person saw.

Dynamic rendering

Serving pre-rendered HTML to crawlers and the client-rendered app to browsers. Two code paths to keep in sync, and the crawler path is the one nobody looks at, so it goes stale quietly.

DOM depth

How deeply nested the document is. Deep trees slow headless rendering and can hit parse or time limits before the content near the bottom is reached.

Edge transformation

Rewriting HTML, headers or status codes at the CDN before the response reaches the client. It can mask an origin misconfiguration from anyone auditing from inside, and hide an edge fault from anyone auditing the origin.

202 interstitial

A homepage answering HTTP 202 Accepted. 202 is an acknowledgement, not a page: a queue or verification step is answering in place of the site, so nothing measured through it describes the site. The scan is refused rather than graded, and the per-identity table shows which named crawler, if any, was handed the real page. Reported as HOMEPAGE_IS_INTERSTITIAL.

rate limiting

A server's answer to a client asking too fast, usually HTTP 429, sometimes a silent slowdown. For measurement the honest consequence is unverifiable: a 429 says the client was throttled and nothing about what it would have received. A scanner that records a throttled fetch as a refusal has invented a finding, and one that retries until it gets through has measured a different moment than the one it reports. The verdict for a throttled path is that it was not measured.

bot management platform

The product category, sometimes called a bot shield, that stands between the origin and every automated client and decides what each one gets. It is the source of challenge page at 200, uniform refusal, and graded through a wall: the shield answers a scanner and most AI crawlers identically, and the site owner may not know it is answering at all, because the owner's own browser passes. What it decides is set by configuration and by the provider's classification, so a scan reports the wall as the wall and does not attribute the response to the site's content.

Measurement and evidence

The vocabulary for saying how strongly something is known.

evidence level

How directly a finding was observed: measured here, inferred from a pattern, or reported by a third party. Stating it is what separates a finding from an assertion.

false positive

A finding that accuses a site of something untrue. On an audit tool this is the most expensive defect class, because it costs the reader trust in every other finding on the page.

denominator note

An explicit statement of what a rate was divided by. A forgery rate that counts unverifiable requests as forged inflates; counting them as genuine understates. Neither is defensible without saying which was done.

percentile rank

Where a value sits within a comparison group. Only meaningful with the group described, since a rank against a seeded corpus is not a rank against the web. measured on /data

population baseline

What share of a measured population does the thing being scored. Without it, a finding says a file is missing and cannot say whether missing is normal or unusual. measured on /data

cohort asymmetry

The gap between how many sites look clean today and how many were broken at least once during a window. A single crawl at any budget can only produce the first number.

self-audit invariant

A check comparing two of a scanner's own outputs to catch contradictions inside its own record. Every one worth encoding is a wrong answer that was actually shipped.

unscored measurement

A finding reported without moving the grade, used where the population is not yet measured or where the applicability of the check depends on the site. Scoring a site for lacking something it has no reason to have is the most common defect in automated audits.

dated series

The same measurement repeated on a schedule and kept. It is the one class of evidence that cannot be produced retroactively at any budget, because it exists only if something was already watching.

change receipt

An artifact proving what moved between two dated measurements, and that neither was edited afterwards. It proves existence, integrity and difference - never that either measurement was correct or that the change caused anything. measured on /pricing

regression

A defect that was fixed and returned. Distinguished from a new finding because it points at a process failure rather than an oversight, and a permanent failure gets noticed while a recurring one only shows up if something was looking at the right moment.

finding age

How long a defect has been present, measured across repeated scans. Converts missing into missing since a date, which is the difference between a snapshot and a record. measured on /pricing

graded through a wall

The self-audit invariant that fires when a record shows the marks of a uniform refusal, or is marked refused, and still carries a grade, a score, a section score or headroom anywhere. A number produced through a wall describes the wall, so the invariant quarantines the report rather than letting it accuse the site. It is raised on the report's own audit, never as a finding against the site.

self-audit

A set of checks a measuring tool runs on its own output after every measurement, comparing two of its own fields for contradictions. A violation is always a bug in the tool, never a finding about the site, and it cannot move a grade. This scanner runs seven such invariants; every one encodes a defect that actually shipped.

placeholder in record

A string from the code that leaked into a stored result: NaN, undefined, [object Object], null inside a URL. A self-audit rule scans every record for them as whole tokens, because a substring match accuses a URL slug that happens to contain the letters.

fetcher disagreement

A self-audit invariant that fires when two layers of the same scanner describe one path differently: one says a file is live, the other says it is not. It is how the scanner learned that an llms.txt answering 200 with HTML was being called present by one layer and absent by another.

own-host loop

A serverless function cannot fetch a route it serves itself; the request never leaves the runtime. For a scanner built that way, any URL on its own host is unreachable from inside a scan. The honest outcome is unverifiable, never dead, and the scanner cannot grade itself from inside.

score version

A number stamped on every stored record naming the scoring rules the numbers were computed under. When thresholds move or sections are added, the version increments; history refuses to draw a delta across the boundary and the corpus percentile refuses to rank across versions, so a grade never changes because the scanner changed while looking like the site did.

row review

Reading every scored row of a rubric as a set, against the pass or fail outcomes of several sites the reviewer understands. A row that fails a site that should pass is measuring the wrong thing. The output is moved thresholds, gated rows, and rows unscored because they counted twice or counted cosmetics.

double-counted defect

One fault scored by two rows: a mechanism row and an outcome row for the same thing, such as a cache status and the fetch time that status determines. The outcome row stays scored; the mechanism row is shown and unscored, or the fault costs a site twice.

cosmetic row

A measurement that cannot affect whether a machine reaches, reads or quotes a page: a touch icon, a theme colour, a favicon format. Shown on a report because it is true, unscored because a grade is a claim about machine readability and a cosmetic row would dilute it.

gated row

A scored row that applies only when the site is the kind of subject the row was written for. Credential edges apply to local and professional services, not to a publisher or a store; expertise properties apply when a Person is declared in JSON-LD, not to an h-card. Outside its gate the row reads unmeasured, never failed.

log file analysis

Reading the server's own request records to learn who fetched what, when, from which address, and what status they received. It is the only first-party evidence of crawler behaviour; everything else is inference from the outside. Every forge rate and verification figure in a report is a statement about logs, whether or not the tool had them, so a rate computed without access to the logs is inferred and should be labelled so. Edge providers keep their own logs, which see requests the origin never receives.

AI referral attribution

Measuring the visits that arrived from an AI answer, the traffic-side complement to share of voice. It is weaker than it looks. Referrer granularity varies by engine, some send none, and agentic browsers send a browser's. A visit proves a person followed a link once; it is not evidence that the page influenced any future answer, and a page cited often with no click-through is invisible to it entirely. Useful as a trend on one site; not comparable across engines or sites.

Proof and provenance

Showing a record existed at a date without asking anyone to trust us.

OpenTimestamps

A scheme that anchors a hash into the Bitcoin blockchain so a document can be proved to have existed at a time. It proves existence and integrity; it proves nothing about whether the content was correct. measured on /proof

anchoring

Committing a hash to a public ledger so it cannot be backdated. Useful for a report whose value depends on the date it was taken. measured on /proof

Content hash

A deterministic digest of a payload. Two parties can confirm they read identical bytes by comparing one short string instead of the whole document.

Manifest tampering

Editing a change log or metadata record after the fact. Hash-chained records make it detectable: the edit no longer matches the digest that was sealed at the time.

C2PA

The Coalition for Content Provenance and Authenticity specification: a signed manifest bound to an image, video or audio file recording how it was created and what was done to it, often surfaced as Content Credentials. It establishes provenance, the chain of edits, and nothing about rights: a manifest does not prove copyright ownership, and an asset with no manifest is not thereby unowned or synthetic. Most assets on the web carry none, so absence is the ordinary case, not a finding.

CAWG training and data mining assertion

An assertion from the Creator Assertions Working Group that travels inside a C2PA manifest and states the creator's preference about AI training and data mining for that specific asset. Because it is signed into the file, it survives copying in a way a robots.txt rule cannot. Attaching it requires signing the asset, which almost no publishing pipeline does, so adoption rather than validity is the limit; it also binds only readers who look inside manifests, which crawlers largely do not.

content-addressed storage

Storage where an object's identifier is derived from a hash of its bytes, as in an IPFS CID, so the address itself proves what it delivers: change one byte and the identifier changes. It is a stronger immutability claim than a change receipt, which proves a difference between two dated observations. It says nothing about authorship or time: anyone can store any bytes, and the identifier does not say when. Paired with a timestamp proof it becomes a dated, verifiable copy; alone it is only a verifiable copy.

What a glossary cannot settle

No engine publishes its ranking function, so every definition here describing influence is a description of a mechanism, not a lever. Terms marked as proposals — llms.txt, agents.md, entity maps — are bets with a low cost, not standards anyone has committed to reading. Where adoption figures exist they are on the dataset page, with the population they were counted against.

Questions about this page

QWhat is an AI visibility score?
A 0–100 measure of whether an answer engine can reach a site, read it and quote a fact about it, computed from scored sections in that order. It measures what a site controls; it does not measure whether an engine actually cites the site.
QWhat is llms.txt?
A plain-text or Markdown file at /llms.txt that names a site, what it does and the pages worth reading, for language-model agents. Adoption is low; the scanner reports it as absent when the path returns HTML or a challenge page with a 200.
QWhat is a verified crawler?
A request whose user-agent names a crawler and whose source IP falls inside the ranges that operator publishes. A user-agent alone is a claim; only the address can confirm it.
QWhy do the definitions state limits?
Because terms like citation, visibility and forged crawler are routinely used to mean more than any measurement supports. Each entry says what evidence the term can rest on.