CrawlCheck

Glossary · Measurement and evaluation

LLM-as-judge

A language model scoring another model’s output. It scales evaluation but imports the judge model’s biases, preferences, and sensitivity to prompt wording.

Measurement and evaluation

How a measurement earns trust: evaluation sets, error rates, calibration, reproducibility, and the biases that make a number look better than it is.

Benchmark · Evaluation set · Golden dataset · Ground truth · Human evaluation · Pairwise evaluation · Pointwise evaluation · Rubric · Inter-rater agreement · Accuracy · F1 score · False positive rate · False negative · True positive · True negative · Calibration · Confidence score · Coverage · Abstention · Reproducibility · Repeatability · Variance · Confidence interval · Sample size · Sampling bias · Selection bias · Survivorship bias · Data leakage · Test contamination

Human evaluation  ·  Pairwise evaluation

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.