Optimising to be quoted inside a generated answer rather than ranked in a list of links. The practices overlap heavily with SEO; what differs is the outcome measured. No engine publishes a ranking function for answers, so any claim to optimise for one is inference from correlation.
A newer name for the same activity as AEO. The two terms are used interchangeably in practice and no consistent technical distinction has been established between them.
Measured: GEO_COUNTRY_MISMATCH 0.1%
A composite number sold by several tools to represent how visible a brand is to language models. There is no shared definition, no published methodology across vendors, and no engine endorses any of them. Two tools can score the same site 40 and 85 and both be internally consistent.
The proportion of sampled answers mentioning your brand versus named competitors. Meaningful only alongside its sample size, its prompt set and its date. A share of voice quoted without those three is a number without a denominator.
The fixed list of questions a monitoring tool sends to models to see who gets mentioned. Results move when the prompt set changes, so a prompt set that is not published cannot be audited or reproduced.
A model naming your brand in an answer. Distinct from a citation: a mention carries no link and no source attribution, so it cannot be traced back to a page you control.
Retrieval-augmented generation. The model retrieves documents at question time and writes an answer grounded in them. Under RAG your page competes to be RETRIEVED, which is a different problem from being in the training data.
Constraining a generated answer to retrieved source material. Grounding reduces invention but does not eliminate it: a grounded answer can still misattribute which source said what.
Text absorbed during model training. Content in training data may influence answers without ever being cited, and cannot be updated or withdrawn once a model is trained. This is why blocking training crawlers and allowing retrieval crawlers is a coherent policy rather than a contradiction.
A numeric vector representing text, used to find passages similar in meaning rather than in wording. Nothing about a page is retrievable by embedding if the page was never fetched in the first place.
Splitting a page into passages before indexing. A page can be retrieved as one passage that reads badly out of context, which is why a self-contained paragraph is easier to quote than one that depends on the paragraph above it.
The amount of text a model can consider at once. A page that exceeds what a retriever will pass along is truncated silently, and the part that survives is not necessarily the part you would have chosen.
The unit a model reads, roughly a word fragment. Token counts matter because markup, navigation and boilerplate consume the same budget as the content they surround.
A confident statement a model generates that is not supported by any source. Distinct from hallucinated attribution, where the claim is right but the credited source is wrong.
An engine declining to answer. A refusal is not evidence about your site: it is evidence about the engine policy at that moment, and treating it as a site defect inverts cause and effect.
Measured: ANSWER_ENGINE_REFUSED 3.5% · ORIGIN_REFUSED_SCANNER 0.2% · HOMEPAGE_REFUSED 0.8%
Google generated answers shown above the results. Appearance is not controllable and not guaranteed, and the sources shown change between identical searches, so absence on one sample proves nothing.
A system that returns a composed answer with citations rather than a list of links. Its retrieval step is what a site can influence; its generation step is not.
Measured: ANSWER_ENGINE_REFUSED 3.5%
A source an answer engine names when composing an answer. A citation proves a document was retrieved and used once; it does not describe the model's ranking and it decays as weighting changes.
The step in which an answer engine selects documents to compose from. Everything a site controls acts here, which is why machine-readable clarity outperforms persuasion.
An answer that credits a claim to a source which does not contain it. The defence is not more content but content that states its own scope, dates and limits.
Running lexical search (BM25 or similar) and dense vector search together at the retrieval step, then merging the results. It exists because pure semantic search misses exact strings — part numbers, model names, a company name that looks like an ordinary word.
A second scoring pass, usually a cross-encoder, that reorders the chunks retrieval returned before the generator sees them. A page can win retrieval and still be dropped here, which is why "we were retrieved" and "we were cited" are different claims.
Scoring individual text segments rather than whole pages. It is why one clearly written paragraph can be cited from a page that is otherwise thin, and why page-level authority metrics predict citation poorly.
An engine rewriting one user prompt into several internal queries before searching. Optimising for the prompt a person typed ignores the queries the system actually ran, and neither we nor any vendor can observe those directly.
The point where text is split to fit a token limit. A split between a premise and its conclusion leaves both halves less usable, and the split is made by the engine, not by you — which is why short self-contained passages survive it better.
The lexical half of hybrid search: a ranking function that scores a document by how often a query term appears in it, discounted by how common the term is across the corpus and by document length. It finds exact strings that embedding similarity blurs, a part number, a company name that is also an ordinary word. It cannot find a paraphrase, which is why retrieval systems pair it with dense retrieval rather than choosing one.
The model class usually behind re-ranking. It reads the query and a candidate passage together, in one pass, and emits a relevance score, which is more accurate than comparing two separately computed vectors and far slower. It runs only on the short list retrieval has already produced. A passage that retrieval never surfaced is never scored by it, so improving how a page reads helps here only after the page is findable at all.
Google's conversational search surface, a separate tab and mode from the AI Overview that appears above ordinary results. Both draw on the Search index, but they fan a query out differently, cite differently, and refresh on different cycles, so a page can be cited in one and absent from the other. We appear in Google's AI results is therefore two claims, and a measurement that does not say which surface it read is not reproducible.
A ranking concept with a Google patent lineage: a document is scored for what it adds beyond the documents the user, or the retrieval set, has already seen. It explains a pattern visible from outside, that a page restating the consensus answer loses to a page contributing a figure, a method, or a dated observation nobody else has. The score itself is not observable from outside an engine, so it is a working explanation, not a measurement, and this scanner does not claim to compute it.
The semantic half of hybrid search: a query and each candidate passage are turned into vectors by an embedding model and matched by similarity, so a paraphrase can be found. Embedding names the vector; dense retrieval names the search step those vectors exist for. It misses exact strings that BM25 finds, it inherits every bias of the model that made the vectors, and a passage that was chunked badly is a bad vector however well it was written.