CrawlCheck

Glossary · area 7 of 19

Measurement and evidence

The vocabulary for saying how strongly something is known.

41 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

41terms in this area
1with a live finding rate

evidence level

How directly a finding was observed: measured here, inferred from a pattern, or reported by a third party. Stating it is what separates a finding from an assertion.

false positive

A finding that accuses a site of something untrue. On an audit tool this is the most expensive defect class, because it costs the reader trust in every other finding on the page.

denominator note

An explicit statement of what a rate was divided by. A forgery rate that counts unverifiable requests as forged inflates; counting them as genuine understates. Neither is defensible without saying which was done.

percentile rank

Where a value sits within a comparison group. Only meaningful with the group described, since a rank against a seeded corpus is not a rank against the web.

population baseline

What share of a measured population does the thing being scored. Without it, a finding says a file is missing and cannot say whether missing is normal or unusual.

cohort asymmetry

The gap between how many sites look clean today and how many were broken at least once during a window. A single crawl at any budget can only produce the first number.

self-audit invariant

A check comparing two of a scanner's own outputs to catch contradictions inside its own record. Every one worth encoding is a wrong answer that was actually shipped.

unscored measurement

A finding reported without moving the grade, used where the population is not yet measured or where the applicability of the check depends on the site. Scoring a site for lacking something it has no reason to have is the most common defect in automated audits.

dated series

The same measurement repeated on a schedule and kept. It is the one class of evidence that cannot be produced retroactively at any budget, because it exists only if something was already watching.

change receipt

An artifact proving what moved between two dated measurements, and that neither was edited afterwards. It proves existence, integrity and difference - never that either measurement was correct or that the change caused anything.

regression

A defect that was fixed and returned. Distinguished from a new finding because it points at a process failure rather than an oversight, and a permanent failure gets noticed while a recurring one only shows up if something was looking at the right moment.

finding age

How long a defect has been present, measured across repeated scans. Converts missing into missing since a date, which is the difference between a snapshot and a record.

graded through a wall

The self-audit invariant that fires when a record shows the marks of a uniform refusal, or is marked refused, and still carries a grade, a score, a section score or headroom anywhere. A number produced through a wall describes the wall, so the invariant quarantines the report rather than letting it accuse the site. It is raised on the report's own audit, never as a finding against the site.

self-audit

A set of checks a measuring tool runs on its own output after every measurement, comparing two of its own fields for contradictions. A violation is always a bug in the tool, never a finding about the site, and it cannot move a grade. This scanner runs seven such invariants; every one encodes a defect that actually shipped.

placeholder in record

A string from the code that leaked into a stored result: NaN, undefined, [object Object], null inside a URL. A self-audit rule scans every record for them as whole tokens, because a substring match accuses a URL slug that happens to contain the letters.

fetcher disagreement

A self-audit invariant that fires when two layers of the same scanner describe one path differently: one says a file is live, the other says it is not. It is how the scanner learned that an llms.txt answering 200 with HTML was being called present by one layer and absent by another.

own-host loop

A serverless function cannot fetch a route it serves itself; the request never leaves the runtime. For a scanner built that way, any URL on its own host is unreachable from inside a scan. The honest outcome is unverifiable, never dead, and the scanner cannot grade itself from inside.

score version

A number stamped on every stored record naming the scoring rules the numbers were computed under. When thresholds move or sections are added, the version increments; history refuses to draw a delta across the boundary and the corpus percentile refuses to rank across versions, so a grade never changes because the scanner changed while looking like the site did.

row review

Reading every scored row of a rubric as a set, against the pass or fail outcomes of several sites the reviewer understands. A row that fails a site that should pass is measuring the wrong thing. The output is moved thresholds, gated rows, and rows unscored because they counted twice or counted cosmetics.

double-counted defect

One fault scored by two rows: a mechanism row and an outcome row for the same thing, such as a cache status and the fetch time that status determines. The outcome row stays scored; the mechanism row is shown and unscored, or the fault costs a site twice.

cosmetic row

A measurement that cannot affect whether a machine reaches, reads or quotes a page: a touch icon, a theme colour, a favicon format. Shown on a report because it is true, unscored because a grade is a claim about machine readability and a cosmetic row would dilute it.

gated row

A scored row that applies only when the site is the kind of subject the row was written for. Credential edges apply to local and professional services, not to a publisher or a store; expertise properties apply when a Person is declared in JSON-LD, not to an h-card. Outside its gate the row reads unmeasured, never failed.

log file analysis

Reading the server's own request records to learn who fetched what, when, from which address, and what status they received. It is the only first-party evidence of crawler behaviour; everything else is inference from the outside. Every forge rate and verification figure in a report is a statement about logs, whether or not the tool had them, so a rate computed without access to the logs is inferred and should be labelled so. Edge providers keep their own logs, which see requests the origin never receives.

AI referral attribution

Measuring the visits that arrived from an AI answer, the traffic-side complement to share of voice. It is weaker than it looks. Referrer granularity varies by engine, some send none, and agentic browsers send a browser's. A visit proves a person followed a link once; it is not evidence that the page influenced any future answer, and a page cited often with no click-through is invisible to it entirely. Useful as a trend on one site; not comparable across engines or sites.

finding fingerprint

A stable identifier for a finding, the code joined to the path it was found on, that stays the same across scans of the same domain. It is what lets a scanner say a finding is new, persisted, worsened, improved, regressed or resolved, rather than reporting each scan as if the last had never happened.

finding state

The lifecycle position of a finding relative to the previous scan: new, persisted, worsened, improved, regressed, resolved, or rule-changed. Resolved means it was present last time and absent now; it does not by itself say the site was fixed, because a rule change or a transient answer can produce the same transition.

headroom

The distance between a section's current score and the top of its optimal range, listed per section so the largest gains are visible. It is a measurement of the score, not of the business: a section with no headroom is at the ceiling of what the scanner measures, which is not the same as having nothing to improve.

fix list

The ranked list of findings with remediation text and the evidence bytes behind each, ordered by what would move the grade most. The free report names every finding and shows the top one in full; the ranked list with every fix is the licensed product. The ranking is by the scanner's weights, which are published.

render gap

The difference between the text a page serves in its HTML and the text that exists only after JavaScript has run, measured as the share of rendered words absent from the delivered bytes. A crawler that does not execute scripts never sees the gap's contents; measuring it needs a real browser, which is what a browser extension provides.

Measured: CONTENT_NEEDS_JAVASCRIPT 0.6%

lab data

Performance measured by loading a page once in a controlled environment, as Lighthouse does. It is repeatable in setup and not in result: the same page varies by tens of percent run to run, so one lab run cannot support a finding. Field data, gathered from real visitors, is what the ranking systems use.

Lighthouse

Google's open-source page-auditing tool, run inside a browser, producing performance, accessibility and best-practice scores from a single simulated load. Its scores are lab data; a single run varies enough that any check built on one number is measuring noise. Useful for diagnosis, not for a grade.

RUM

Real user monitoring: performance measured in visitors' own browsers and reported back, giving field data for pages and sites the public datasets never cover. A small site gets no Chrome UX Report entry at all, so a first-party beacon is the only route to the metrics that rankings actually read.

LCP

Largest Contentful Paint, the time until the largest visible element has rendered, one of the Core Web Vitals. Under 2.5 seconds is the threshold. It is dominated by the hero image or heading and by whatever blocks it; a lazy-loaded hero image is the most common self-inflicted cause.

CLS

Cumulative Layout Shift, a score for how much visible content moves after it first renders, one of the Core Web Vitals. Under 0.1 is the threshold. Images without declared dimensions, late-loading fonts and injected banners are the usual causes; it is measured across the whole visit, not at load.

INP

Interaction to Next Paint, the delay between a user's input and the next frame the page paints, one of the Core Web Vitals since 2024. Under 200 milliseconds is the threshold. It is a field metric by nature: a lab run with no interactions cannot measure it at all.

topicality

How closely a page's entities and vocabulary match the subject an engine believes it is about, as distinct from whether the page is well-formed. A page can pass every technical check and still not be about what its title claims; topicality is the check that reads the words rather than the markup.

intent coverage

Whether a site has a page for each of the questions its audience asks about its subject: how much, how long, is it worth it, what does it include, near me. Measured by comparing declared services and areas against the pages that exist; a service with no page has no answer for the engine to quote.

canary

A known input sent through a system on a schedule so that its arrival, not the absence of an error, proves the path works. A canary message to an alert address, confirmed by reading the mailbox, catches a suppressed or bounced destination that every send-side check reports as success.

citation rate

The share of answer-engine responses on a set of prompts in which a site is cited as a source. It depends entirely on the prompt set, the engine and the day, so a rate without those three stated is decoration. It is not the same as a mention, which names the brand without linking it.

presence rate

The share of runs of the same prompt, across models and days, in which a page or brand appears in the answer at all. It measures stability rather than rank: a page present in nine of ten runs is a fact an engine has settled on, a page present in three is not.

answer stability

How much an engine's answer to the same prompt changes between runs, models and days. High churn means a citation observed once is weak evidence; measuring it needs repeated runs, which is why a single screenshot of an answer proves almost nothing about visibility.

← Delivery and rendering  ·  Proof and provenance →

All 668 terms across 19 areas.