Glossary · Models and generation
Context caching
Reusing previously processed prompt prefixes to reduce repeated inference work. It lowers latency or cost but does not enlarge the model’s context window or refresh cached facts.
Terms this definition uses
Models and generation
How a model produces text, what its training and decoding controls change, and what none of them can prove about a source.
Autoregressive model · Transformer · Attention · Self-attention · Parameter · Inference · Pretraining · Fine-tuning · Supervised fine-tuning · Instruction tuning · RLHF · Preference optimization · Knowledge distillation · Quantization · Mixture of experts · Multimodal model · Vision-language model · Small language model · Tokenizer · Temperature · Top-p sampling · Deterministic decoding · System prompt · Developer message · Tool call · Function calling · Structured output · Constrained decoding · Chain of thought · Reasoning trace · Context compaction · Model routing · Model drift · Model version
← Reasoning trace · Context compaction →
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.