CrawlCheck

Glossary · Models and generation

RLHF

Reinforcement learning from human feedback, where preference judgments shape model behavior. It aligns outputs with sampled raters and policies, not with an objective definition of truth.

Terms this definition uses

Object

Models and generation

How a model produces text, what its training and decoding controls change, and what none of them can prove about a source.

Autoregressive model · Transformer · Attention · Self-attention · Parameter · Inference · Pretraining · Fine-tuning · Supervised fine-tuning · Instruction tuning · Preference optimization · Knowledge distillation · Quantization · Mixture of experts · Multimodal model · Vision-language model · Small language model · Tokenizer · Temperature · Top-p sampling · Deterministic decoding · System prompt · Developer message · Tool call · Function calling · Structured output · Constrained decoding · Chain of thought · Reasoning trace · Context caching · Context compaction · Model routing · Model drift · Model version

Instruction tuning  ·  Preference optimization

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.