Glossary · Models and generation
Mixture of experts
A model architecture that routes each input through selected parameter groups rather than activating the whole model. The advertised total parameter count therefore overstates the compute used per request.
Terms this definition uses
Models and generation
How a model produces text, what its training and decoding controls change, and what none of them can prove about a source.
Autoregressive model · Transformer · Attention · Self-attention · Parameter · Inference · Pretraining · Fine-tuning · Supervised fine-tuning · Instruction tuning · RLHF · Preference optimization · Knowledge distillation · Quantization · Multimodal model · Vision-language model · Small language model · Tokenizer · Temperature · Top-p sampling · Deterministic decoding · System prompt · Developer message · Tool call · Function calling · Structured output · Constrained decoding · Chain of thought · Reasoning trace · Context caching · Context compaction · Model routing · Model drift · Model version
← Quantization · Multimodal model →
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.