CrawlCheck

Glossary · area 9 of 19

Models and generation

How a model produces text, what its training and decoding controls change, and what none of them can prove about a source.

35 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

35terms in this area
0with a live finding rate

Autoregressive model

A model that generates each token from the tokens preceding it. It predicts a continuation, not whether that continuation is factual, current, or supported by a source.

Transformer

A neural-network architecture built around attention mechanisms for processing relationships across a sequence. The architecture explains how information is processed, not what data a deployed model learned.

Attention

A mechanism that weights which input representations matter while computing another representation. A high internal weight is not a citation, explanation, or human-readable measure of causal importance.

Self-attention

Attention computed among positions within the same sequence. It helps a model relate distant words, but it does not remove context limits or guarantee coherent long-document reasoning.

Parameter

A learned numerical value adjusted during training. Parameter count describes model scale, not accuracy, knowledge coverage, reasoning quality, or the amount of copyrighted material represented.

Inference

Running a trained model on an input to produce an output. It uses learned parameters and supplied context; it does not update the model unless a separate learning process occurs.

Pretraining

Large-scale learning from broad data before adaptation to a particular task. It creates general capabilities but cannot keep facts current after the training cutoff.

Fine-tuning

Additional training that adapts an existing model using a narrower dataset. It can change behavior or domain performance, but it cannot guarantee compliance with every prompt.

Supervised fine-tuning

Fine-tuning on examples pairing inputs with desired outputs. It teaches demonstrated patterns and may fail when production requests differ materially from the examples.

Instruction tuning

Training on tasks expressed as natural-language instructions. It improves instruction following but does not establish that the model understands an instruction as a person would.

RLHF

Reinforcement learning from human feedback, where preference judgments shape model behavior. It aligns outputs with sampled raters and policies, not with an objective definition of truth.

Preference optimization

Training that favors outputs ranked above alternatives by humans or another model. It changes response tendencies; it does not make the preferred response factually correct.

Knowledge distillation

Training a smaller model to reproduce behavior from a larger teacher. It can reduce cost, but omissions and errors from the teacher can also be transferred.

Quantization

Representing model weights or activations with lower numerical precision. It reduces memory and compute requirements, sometimes at the cost of output quality or stability.

Mixture of experts

A model architecture that routes each input through selected parameter groups rather than activating the whole model. The advertised total parameter count therefore overstates the compute used per request.

Multimodal model

A model that accepts or produces more than one medium, such as text, images, audio, or video. Supported input does not mean equal competence across modalities.

Vision-language model

A multimodal model designed to relate images and language. It can describe visible patterns but can misread small text, spatial relationships, or image provenance.

Small language model

A language model optimized around fewer parameters or lower resource requirements. “Small” has no universal threshold and does not by itself mean local, private, or weak.

Tokenizer

Software that converts content into the token units a model processes. Different tokenizers split identical text differently, so token counts are model-specific rather than properties of the page.

Temperature

A decoding control that changes how strongly generation favors the highest-probability next tokens. Lower values reduce variation but do not make an answer deterministic across infrastructure or model versions.

Top-p sampling

Generation restricted to the smallest token set whose cumulative probability reaches a threshold. It controls diversity, not factuality or source selection.

Deterministic decoding

Selecting outputs without intentional sampling, commonly by choosing the highest-scoring next token. The same request can still change after model, prompt, tool, or serving updates.

System prompt

High-priority instructions supplied by the application operating a model. It shapes behavior but is not an impenetrable security boundary against conflicting or injected content.

Developer message

Application-level instructions placed below system policy and above ordinary user content in some model APIs. Its exact priority and availability depend on the provider.

Tool call

A structured request from a model asking application code to invoke an external function. The application, not the model, decides whether and how the action executes.

Function calling

Constraining a model to select named functions and produce arguments matching declared interfaces. Valid arguments can still request an unsafe, unauthorized, or factually mistaken action.

Structured output

Model output constrained to a declared shape such as JSON Schema. Structural validity says nothing about whether the field values are true or complete.

Constrained decoding

Restricting token generation so outputs follow a grammar, schema, or allowed vocabulary. It prevents some formatting failures, not semantic errors inside the permitted structure.

Chain of thought

Intermediate natural-language reasoning produced or represented during problem solving. A plausible chain can rationalize a wrong answer and should not be treated as proof of the actual internal computation.

Reasoning trace

A recorded sequence of intermediate model or agent steps. It supports debugging when authentic and complete, but providers may summarize, hide, or generate traces separately from hidden computation.

Context caching

Reusing previously processed prompt prefixes to reduce repeated inference work. It lowers latency or cost but does not enlarge the model’s context window or refresh cached facts.

Context compaction

Replacing older conversation material with a shorter retained representation. It extends long-running work while risking the loss of qualifications, provenance, and low-salience facts.

Model routing

Choosing among models based on cost, difficulty, latency, policy, or modality. A routed product name does not identify which model answered unless that decision is logged.

Model drift

Output behavior changing over time because weights, prompts, tools, safety rules, retrieval, or serving infrastructure changed. A changed answer does not identify which layer drifted.

Model version

A named or dated release of model behavior. A stable marketing name may point to updated weights, so reproducibility requires the exact version or snapshot identifier where available.

← Proof and provenance  ·  Retrieval engineering →

All 668 terms across 19 areas.