CrawlCheck

Glossary · area 15 of 19

Measurement and evaluation

How a measurement earns trust: evaluation sets, error rates, calibration, reproducibility, and the biases that make a number look better than it is.

30 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

30terms in this area
0with a live finding rate

Benchmark

A fixed task set and scoring procedure used to compare systems. Performance can be optimized to the benchmark and may not transfer to production work.

Evaluation set

Inputs and expected judgments reserved for measuring a system. Leakage into training, prompting, or tuning makes the score overstate generalization.

Golden dataset

A curated set treated as the reference truth for evaluation. It is only as reliable as its coverage, labels, update process, and disagreement handling.

Ground truth

The reference answer or label against which output is evaluated. In subjective or changing domains, it may be a negotiated judgment rather than objective truth.

Human evaluation

People judging outputs against stated criteria. It captures qualities automated metrics miss while introducing rater variation, cost, and fatigue.

LLM-as-judge

A language model scoring another model’s output. It scales evaluation but imports the judge model’s biases, preferences, and sensitivity to prompt wording.

Pairwise evaluation

Comparing two outputs and choosing the better one. It is often easier than absolute scoring but does not show whether either output meets a minimum standard.

Pointwise evaluation

Scoring one output independently against a rubric. It supports thresholds while making calibration differences between judges more visible.

Rubric

A defined set of criteria and scoring levels for judgments. A detailed rubric improves consistency but cannot correct missing or biased criteria.

Inter-rater agreement

The extent to which independent reviewers assign compatible judgments. High agreement can reflect clear criteria or shared bias; low agreement signals ambiguity that averages conceal.

Accuracy

The share of evaluated predictions judged correct. It can mislead on imbalanced tasks and says nothing about the severity of errors.

F1 score

The harmonic mean of precision and recall. It balances the two under one threshold while hiding which component is limiting performance.

False positive rate

The share of actually negative cases incorrectly labeled positive. It requires a trustworthy negative denominator and should not be confused with false discovery rate.

False negative

A relevant condition the system fails to detect. It is measurable only where the condition’s presence is independently known.

True positive

A detected condition confirmed as present by the reference standard. It measures agreement with that standard, not necessarily real-world utility.

True negative

A correctly rejected or absent condition according to the reference standard. Large easy-negative populations can make overall accuracy look stronger than useful performance.

Calibration

Agreement between predicted confidence and observed correctness over comparable cases. A calibrated 80 percent confidence should be correct about 80 percent of the time, not every time.

Confidence score

A numerical estimate associated with a prediction or classification. Unless calibrated and scoped, it is not a literal probability that the answer is true.

Coverage

The share of an intended population for which the system produces a measurable result. Excluding failures can improve reported quality while reducing actual coverage.

Abstention

A system declining to classify or answer when evidence is insufficient. It can reduce false claims while lowering coverage and user satisfaction.

Reproducibility

The ability to obtain materially consistent findings from the documented inputs, versions, configuration, and procedure. Model and web drift can prevent exact repetition despite honest methods.

Repeatability

Consistency when the same operator repeats a measurement under the same conditions. It is narrower than reproducibility across teams or environments.

Variance

The spread of repeated observations around their average. A single result cannot reveal variance, and a low average can coexist with unstable outliers.

Confidence interval

A range produced by a stated statistical procedure to quantify uncertainty around an estimate. It does not mean the fixed true value has that probability of lying inside this completed interval.

Sample size

The number of independent units contributing to an estimate. Large counts do not cure biased selection, duplication, or non-independence.

Sampling bias

Systematic difference between measured cases and the population a claim describes. More observations from the same biased source make the estimate more precise, not more representative.

Selection bias

Distortion caused by which sites, prompts, pages, or users enter analysis. Opt-in datasets especially should not be described as the whole web.

Survivorship bias

Drawing conclusions only from entities still present or successfully measured. Failed, blocked, removed, or abandoned cases can reverse the apparent pattern.

Data leakage

Information from evaluation cases entering training, prompting, retrieval, or tuning. It creates performance that may disappear on genuinely unseen inputs.

Test contamination

Benchmark content appearing in a model’s training or accessible context. High scores then may measure memorization rather than general capability.

← Content architecture and SEO  ·  Agents and protocols →

All 668 terms across 19 areas.