Glossary · Measurement and evaluation
Evaluation set
Inputs and expected judgments reserved for measuring a system. Leakage into training, prompting, or tuning makes the score overstate generalization.
Measurement and evaluation
How a measurement earns trust: evaluation sets, error rates, calibration, reproducibility, and the biases that make a number look better than it is.
Benchmark · Golden dataset · Ground truth · Human evaluation · LLM-as-judge · Pairwise evaluation · Pointwise evaluation · Rubric · Inter-rater agreement · Accuracy · F1 score · False positive rate · False negative · True positive · True negative · Calibration · Confidence score · Coverage · Abstention · Reproducibility · Repeatability · Variance · Confidence interval · Sample size · Sampling bias · Selection bias · Survivorship bias · Data leakage · Test contamination
← Benchmark · Golden dataset →
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.