CrawlCheck

Glossary · area 10 of 19

Retrieval engineering

How pages become chunks, vectors and candidates inside an answer engine, and why being retrieved is not being cited.

45 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

45terms in this area
0with a live finding rate

Corpus

The collection of documents available to a search or retrieval system. It defines what can be found; no ranking method can retrieve a document absent from the corpus.

Document ingestion

Importing source material into a retrieval pipeline. A successful import proves transfer, not that the content parsed correctly or entered the searchable index.

Data connector

An integration that reads from a source such as a website, drive, database, or API. Connector access does not guarantee complete synchronization or permission to reuse the data.

Parsing

Converting source bytes into a structured representation that downstream systems can process. A parser can succeed syntactically while dropping tables, relationships, labels, or reading order.

OCR

Optical character recognition, which converts text visible in images or scans into characters. It enables search but introduces errors that can become confidently repeated facts.

Layout-aware parsing

Extraction that retains headings, columns, tables, captions, and spatial relationships. It reduces structural loss but cannot infer relationships the document never makes explicit.

Semantic chunking

Splitting content at inferred topic or discourse boundaries. It may produce more coherent passages, but the boundaries depend on the model and configuration used.

Fixed-size chunking

Dividing content by a fixed character or token count. It is reproducible and cheap, but it can separate a claim from its evidence or qualification.

Chunk overlap

Repeating material across adjacent chunks so boundary content survives retrieval. It improves continuity while increasing index size and duplicate candidates.

Parent-child retrieval

Indexing small child passages while returning a larger parent document around the match. It balances match precision and context, but the parent may exceed the generator’s usable budget.

Metadata filtering

Restricting retrieval by fields such as date, language, author, or product. It improves precision only when metadata is populated consistently and correctly.

Vector database

A system optimized to store embeddings and retrieve nearby vectors. It finds mathematical similarity, not necessarily factual relevance, authority, permission, or freshness.

Vector index

A data structure that accelerates nearest-neighbor lookup over embeddings. Its configuration trades recall, latency, memory, and update cost rather than maximizing all four.

Sparse vector

A high-dimensional representation in which most values are zero, often preserving lexical features. It supports exact terminology but does not inherently understand paraphrases.

Dense vector

A compact representation in which most dimensions carry values learned by an embedding model. It supports semantic matching but can blur rare names and exact identifiers.

Cosine similarity

A comparison of vector direction that largely ignores magnitude. A high score means geometric closeness under one embedding model, not that two passages state the same fact.

Dot-product similarity

A vector comparison affected by both direction and magnitude. Scores are not comparable across models or indexes without knowing how embeddings were produced and normalized.

Approximate nearest neighbor

Fast retrieval that searches likely nearby vectors instead of exhaustively comparing every record. Its speed comes from accepting that some true neighbors may be missed.

HNSW

Hierarchical Navigable Small World, a graph-based approximate-nearest-neighbor index. Its tuning affects memory, build time, latency, and recall; the acronym is not a quality guarantee.

Top-k

The number of highest-ranked candidates retained at a retrieval or reranking stage. Raising it increases coverage and context cost, not necessarily answer quality.

Retrieval recall

The share of relevant items that a retriever returns within a defined cutoff. It requires a judged relevant set; without one, recall cannot be measured.

Retrieval precision

The share of returned items judged relevant. High precision can coexist with missing many useful sources, so it must be read alongside recall.

Mean reciprocal rank

The average reciprocal position of the first relevant result across queries. It rewards finding one relevant item early and ignores later relevant results.

Normalized discounted cumulative gain

A ranked-retrieval metric that gives more credit to highly relevant results near the top. Its value depends entirely on the relevance judgments and cutoff used.

Relevance score

A system-specific number estimating how well a candidate matches a query. It is usually meaningful only within the same model, index, query, and scoring configuration.

Candidate generation

The fast first stage that retrieves a manageable set for more expensive scoring. Anything excluded here cannot be recovered by later reranking.

Query rewriting

Transforming a request into a form expected to retrieve better results. It can clarify intent or silently change it, so the rewritten query should be observable in auditable systems.

Query expansion

Adding synonyms, entities, or related expressions to broaden matching. It improves recall while risking topic drift and lower precision.

HyDE

Hypothetical Document Embeddings, where a model drafts an imagined answer and retrieves documents similar to it. If the hypothetical answer is wrong, retrieval can be steered toward confirming material.

Multi-query retrieval

Running several reformulations of one request and merging their results. It covers more interpretations at higher latency and can amplify duplicates.

Reciprocal rank fusion

Combining ranked lists by adding scores based on each item’s position. It works without comparable raw scores but does not judge source quality by itself.

Late interaction

Comparing query and document token representations near retrieval time rather than collapsing each into one vector. It can improve matching at greater storage and compute cost.

Lexical retrieval

Matching text through words, stems, or character features. It excels at exact strings and fails when relevant material uses vocabulary the query never shares.

Exact-match retrieval

Requiring a literal identifier, phrase, or field value. It is reliable for known strings and useless when spelling, formatting, or terminology varies.

Freshness boosting

Increasing ranking weight for recently published or updated material. It favors recency, which is not the same as accuracy or authority.

Authority weighting

Adjusting retrieval scores using signals about source reputation or trust. The result inherits whatever assumptions and blind spots define “authority.”

Source diversity

Selecting results across different domains, publishers, or viewpoints. Diversity reduces repetition but can lower average relevance if enforced mechanically.

Deduplication

Removing records judged to represent the same content or entity. Aggressive rules can collapse distinct versions; weak rules leave repeated evidence that falsely appears independent.

Near-duplicate detection

Identifying substantially similar content despite small wording or formatting changes. Similarity thresholds determine whether syndication, updates, and legitimate variants are incorrectly merged.

Semantic cache

Reusing an earlier result when a new request is judged meaningfully similar. It reduces latency but can serve stale or inapplicable answers to superficially related questions.

Retrieval latency

Time spent locating and scoring source material before generation. Low latency does not indicate high recall, relevance, or freshness.

Retrieval cutoff

The rank, score, or count threshold below which candidates are discarded. It makes the pipeline tractable and creates an invisible boundary between considered and unseen sources.

Citation mapping

Linking answer claims or spans to retrieved source passages. It supports auditability only when the mapping reflects actual evidence rather than post-hoc similarity.

Answer synthesis

Combining retrieved material into a composed response. A fluent synthesis can merge incompatible sources or remove their qualifications, even when retrieval was correct.

← Models and generation  ·  Crawling and indexing →

All 668 terms across 19 areas.