CrawlCheck

Glossary · Measurement and evaluation

Evaluation set

Inputs and expected judgments reserved for measuring a system. Leakage into training, prompting, or tuning makes the score overstate generalization.

Measurement and evaluation

How a measurement earns trust: evaluation sets, error rates, calibration, reproducibility, and the biases that make a number look better than it is.

Benchmark · Golden dataset · Ground truth · Human evaluation · LLM-as-judge · Pairwise evaluation · Pointwise evaluation · Rubric · Inter-rater agreement · Accuracy · F1 score · False positive rate · False negative · True positive · True negative · Calibration · Confidence score · Coverage · Abstention · Reproducibility · Repeatability · Variance · Confidence interval · Sample size · Sampling bias · Selection bias · Survivorship bias · Data leakage · Test contamination

Benchmark  ·  Golden dataset

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.