Benchmark
A fixed task set and scoring procedure used to compare systems. Performance can be optimized to the benchmark and may not transfer to production work.
Glossary · area 15 of 19
How a measurement earns trust: evaluation sets, error rates, calibration, reproducibility, and the biases that make a number look better than it is.
30 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
A fixed task set and scoring procedure used to compare systems. Performance can be optimized to the benchmark and may not transfer to production work.
Inputs and expected judgments reserved for measuring a system. Leakage into training, prompting, or tuning makes the score overstate generalization.
A curated set treated as the reference truth for evaluation. It is only as reliable as its coverage, labels, update process, and disagreement handling.
The reference answer or label against which output is evaluated. In subjective or changing domains, it may be a negotiated judgment rather than objective truth.
People judging outputs against stated criteria. It captures qualities automated metrics miss while introducing rater variation, cost, and fatigue.
A language model scoring another model’s output. It scales evaluation but imports the judge model’s biases, preferences, and sensitivity to prompt wording.
Comparing two outputs and choosing the better one. It is often easier than absolute scoring but does not show whether either output meets a minimum standard.
Scoring one output independently against a rubric. It supports thresholds while making calibration differences between judges more visible.
A defined set of criteria and scoring levels for judgments. A detailed rubric improves consistency but cannot correct missing or biased criteria.
The extent to which independent reviewers assign compatible judgments. High agreement can reflect clear criteria or shared bias; low agreement signals ambiguity that averages conceal.
The share of evaluated predictions judged correct. It can mislead on imbalanced tasks and says nothing about the severity of errors.
The harmonic mean of precision and recall. It balances the two under one threshold while hiding which component is limiting performance.
The share of actually negative cases incorrectly labeled positive. It requires a trustworthy negative denominator and should not be confused with false discovery rate.
A relevant condition the system fails to detect. It is measurable only where the condition’s presence is independently known.
A detected condition confirmed as present by the reference standard. It measures agreement with that standard, not necessarily real-world utility.
A correctly rejected or absent condition according to the reference standard. Large easy-negative populations can make overall accuracy look stronger than useful performance.
Agreement between predicted confidence and observed correctness over comparable cases. A calibrated 80 percent confidence should be correct about 80 percent of the time, not every time.
A numerical estimate associated with a prediction or classification. Unless calibrated and scoped, it is not a literal probability that the answer is true.
The share of an intended population for which the system produces a measurable result. Excluding failures can improve reported quality while reducing actual coverage.
A system declining to classify or answer when evidence is insufficient. It can reduce false claims while lowering coverage and user satisfaction.
The ability to obtain materially consistent findings from the documented inputs, versions, configuration, and procedure. Model and web drift can prevent exact repetition despite honest methods.
Consistency when the same operator repeats a measurement under the same conditions. It is narrower than reproducibility across teams or environments.
The spread of repeated observations around their average. A single result cannot reveal variance, and a low average can coexist with unstable outliers.
A range produced by a stated statistical procedure to quantify uncertainty around an estimate. It does not mean the fixed true value has that probability of lying inside this completed interval.
The number of independent units contributing to an estimate. Large counts do not cure biased selection, duplication, or non-independence.
Systematic difference between measured cases and the population a claim describes. More observations from the same biased source make the estimate more precise, not more representative.
Distortion caused by which sites, prompts, pages, or users enter analysis. Opt-in datasets especially should not be described as the whole web.
Drawing conclusions only from entities still present or successfully measured. Failed, blocked, removed, or abandoned cases can reverse the apparent pattern.
Information from evaluation cases entering training, prompting, retrieval, or tuning. It creates performance that may disappear on genuinely unseen inputs.
Benchmark content appearing in a model’s training or accessible context. High scores then may measure memorization rather than general capability.