CrawlCheck

Glossary · Retrieval engineering

Near-duplicate detection

Identifying substantially similar content despite small wording or formatting changes. Similarity thresholds determine whether syndication, updates, and legitimate variants are incorrectly merged.

Terms this definition uses

ARIA

Retrieval engineering

How pages become chunks, vectors and candidates inside an answer engine, and why being retrieved is not being cited.

Corpus · Document ingestion · Data connector · Parsing · OCR · Layout-aware parsing · Semantic chunking · Fixed-size chunking · Chunk overlap · Parent-child retrieval · Metadata filtering · Vector database · Vector index · Sparse vector · Dense vector · Cosine similarity · Dot-product similarity · Approximate nearest neighbor · HNSW · Top-k · Retrieval recall · Retrieval precision · Mean reciprocal rank · Normalized discounted cumulative gain · Relevance score · Candidate generation · Query rewriting · Query expansion · HyDE · Multi-query retrieval · Reciprocal rank fusion · Late interaction · Lexical retrieval · Semantic search · Exact-match retrieval · Freshness boosting · Authority weighting · Source diversity · Deduplication · Semantic cache · Retrieval latency · Retrieval cutoff · Citation mapping · Answer synthesis

Deduplication  ·  Semantic cache

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.