CrawlCheck

Glossary · Models and generation

Context caching

Reusing previously processed prompt prefixes to reduce repeated inference work. It lowers latency or cost but does not enlarge the model’s context window or refresh cached facts.

Terms this definition uses

Inference · context window

Models and generation

How a model produces text, what its training and decoding controls change, and what none of them can prove about a source.

Autoregressive model · Transformer · Attention · Self-attention · Parameter · Inference · Pretraining · Fine-tuning · Supervised fine-tuning · Instruction tuning · RLHF · Preference optimization · Knowledge distillation · Quantization · Mixture of experts · Multimodal model · Vision-language model · Small language model · Tokenizer · Temperature · Top-p sampling · Deterministic decoding · System prompt · Developer message · Tool call · Function calling · Structured output · Constrained decoding · Chain of thought · Reasoning trace · Context compaction · Model routing · Model drift · Model version

Reasoning trace  ·  Context compaction

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.