Corpus
The collection of documents available to a search or retrieval system. It defines what can be found; no ranking method can retrieve a document absent from the corpus.
Glossary · area 10 of 19
How pages become chunks, vectors and candidates inside an answer engine, and why being retrieved is not being cited.
45 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
The collection of documents available to a search or retrieval system. It defines what can be found; no ranking method can retrieve a document absent from the corpus.
Importing source material into a retrieval pipeline. A successful import proves transfer, not that the content parsed correctly or entered the searchable index.
An integration that reads from a source such as a website, drive, database, or API. Connector access does not guarantee complete synchronization or permission to reuse the data.
Converting source bytes into a structured representation that downstream systems can process. A parser can succeed syntactically while dropping tables, relationships, labels, or reading order.
Optical character recognition, which converts text visible in images or scans into characters. It enables search but introduces errors that can become confidently repeated facts.
Extraction that retains headings, columns, tables, captions, and spatial relationships. It reduces structural loss but cannot infer relationships the document never makes explicit.
Splitting content at inferred topic or discourse boundaries. It may produce more coherent passages, but the boundaries depend on the model and configuration used.
Dividing content by a fixed character or token count. It is reproducible and cheap, but it can separate a claim from its evidence or qualification.
Repeating material across adjacent chunks so boundary content survives retrieval. It improves continuity while increasing index size and duplicate candidates.
Indexing small child passages while returning a larger parent document around the match. It balances match precision and context, but the parent may exceed the generator’s usable budget.
Restricting retrieval by fields such as date, language, author, or product. It improves precision only when metadata is populated consistently and correctly.
A system optimized to store embeddings and retrieve nearby vectors. It finds mathematical similarity, not necessarily factual relevance, authority, permission, or freshness.
A data structure that accelerates nearest-neighbor lookup over embeddings. Its configuration trades recall, latency, memory, and update cost rather than maximizing all four.
A high-dimensional representation in which most values are zero, often preserving lexical features. It supports exact terminology but does not inherently understand paraphrases.
A compact representation in which most dimensions carry values learned by an embedding model. It supports semantic matching but can blur rare names and exact identifiers.
A comparison of vector direction that largely ignores magnitude. A high score means geometric closeness under one embedding model, not that two passages state the same fact.
A vector comparison affected by both direction and magnitude. Scores are not comparable across models or indexes without knowing how embeddings were produced and normalized.
Fast retrieval that searches likely nearby vectors instead of exhaustively comparing every record. Its speed comes from accepting that some true neighbors may be missed.
Hierarchical Navigable Small World, a graph-based approximate-nearest-neighbor index. Its tuning affects memory, build time, latency, and recall; the acronym is not a quality guarantee.
The number of highest-ranked candidates retained at a retrieval or reranking stage. Raising it increases coverage and context cost, not necessarily answer quality.
The share of relevant items that a retriever returns within a defined cutoff. It requires a judged relevant set; without one, recall cannot be measured.
The share of returned items judged relevant. High precision can coexist with missing many useful sources, so it must be read alongside recall.
The average reciprocal position of the first relevant result across queries. It rewards finding one relevant item early and ignores later relevant results.
A ranked-retrieval metric that gives more credit to highly relevant results near the top. Its value depends entirely on the relevance judgments and cutoff used.
A system-specific number estimating how well a candidate matches a query. It is usually meaningful only within the same model, index, query, and scoring configuration.
The fast first stage that retrieves a manageable set for more expensive scoring. Anything excluded here cannot be recovered by later reranking.
Transforming a request into a form expected to retrieve better results. It can clarify intent or silently change it, so the rewritten query should be observable in auditable systems.
Adding synonyms, entities, or related expressions to broaden matching. It improves recall while risking topic drift and lower precision.
Hypothetical Document Embeddings, where a model drafts an imagined answer and retrieves documents similar to it. If the hypothetical answer is wrong, retrieval can be steered toward confirming material.
Running several reformulations of one request and merging their results. It covers more interpretations at higher latency and can amplify duplicates.
Combining ranked lists by adding scores based on each item’s position. It works without comparable raw scores but does not judge source quality by itself.
Comparing query and document token representations near retrieval time rather than collapsing each into one vector. It can improve matching at greater storage and compute cost.
Matching text through words, stems, or character features. It excels at exact strings and fails when relevant material uses vocabulary the query never shares.
Retrieval based on learned meaning representations rather than exact wording alone. It finds paraphrases but can return conceptually similar passages that do not answer the question.
Requiring a literal identifier, phrase, or field value. It is reliable for known strings and useless when spelling, formatting, or terminology varies.
Increasing ranking weight for recently published or updated material. It favors recency, which is not the same as accuracy or authority.
Adjusting retrieval scores using signals about source reputation or trust. The result inherits whatever assumptions and blind spots define “authority.”
Selecting results across different domains, publishers, or viewpoints. Diversity reduces repetition but can lower average relevance if enforced mechanically.
Removing records judged to represent the same content or entity. Aggressive rules can collapse distinct versions; weak rules leave repeated evidence that falsely appears independent.
Identifying substantially similar content despite small wording or formatting changes. Similarity thresholds determine whether syndication, updates, and legitimate variants are incorrectly merged.
Reusing an earlier result when a new request is judged meaningfully similar. It reduces latency but can serve stale or inapplicable answers to superficially related questions.
Time spent locating and scoring source material before generation. Low latency does not indicate high recall, relevance, or freshness.
The rank, score, or count threshold below which candidates are discarded. It makes the pipeline tractable and creates an invisible boundary between considered and unseen sources.
Linking answer claims or spans to retrieved source passages. It supports auditability only when the mapping reflects actual evidence rather than post-hoc similarity.
Combining retrieved material into a composed response. A fluent synthesis can merge incompatible sources or remove their qualifications, even when retrieval was correct.