CrawlCheck

Findings · 2026-08-15

Your site was clean the day you audited it. The other days are the finding.

Measuring the same sites every day produces a number a one-off audit structurally cannot: how many were broken at least once. The gap between “broken today” and “ever broken” is the case for a series over a snapshot.

Every audit report answers one question: what was true at the moment we looked. The defects that make that answer misleading are the intermittent ones — a cache that serves stale robots rules for an hour a day, a deploy that briefly drops a header, a rate limiter that refuses crawlers only under load. Check on Tuesday and the site is clean. It was broken on Sunday, and it will be broken on Thursday, and no Tuesday audit at any budget can see either.

,

The measurement

,

This instrument scans a fixed panel of sites every day and keeps each day’s measurements permanently. That makes two percentages computable over the same sites and the same window: how many carry a defect on their latest reading, and how many carried one on at least one day. A single crawl — anyone’s, including ours — can only ever produce the first number. The second needs the sites to have been measured before the question was asked, which is why nobody quoting a one-off audit can produce it.

,

What a single crawl cannot see

64.6% of measured sites carry a defect today. 91.7% carried one at least once between 2026-08-14 and 2026-08-15. 13 look clean now and did not for at least one day in that window.

Across 48 sites measured every dayShare
Carry a defect today64.6%
Carried one at least once in the window91.7%
The gap27.1 points
Look clean now and did not, at least once13

Both numbers describe the same sites over 2 days, 2026-08-14 to 2026-08-15. A one-off crawl at any budget can only ever produce the first row — the second needs the same sites to have been measured before the question was asked. Sites added after the window opened are excluded (6) because they have had fewer chances to be seen broken. No scanned domain is named here or anywhere else on this site.

,

Citing this figure. The numbers above are recomputed from the live record when this page loads, so link the anchor rather than freezing a copy: https://crawlcheck.io/blog/your-site-was-clean-the-day-you-audited-it#the-gap. Sentence form, as of 2026-08-15: Among the 48 sites CrawlCheck measured daily between 2026-08-14 and 2026-08-15, 64.6% carried a machine-layer defect on the latest reading, while 91.7% carried one at least once — a gap of 27.1 points invisible to any single-day audit.

,

The exclusions, because they move the number

,

A site added after the window opened has had fewer chances to be seen broken, so late joiners are excluded from both percentages rather than allowed to drag the “ever” figure down. A site whose latest reading carries no findings count is excluded from both denominators and counted openly. And the whole table refuses to render below its own publishability gate — with too few days, “at least once” and “right now” are the same set by construction, and publishing them as different numbers would be theatre.

,

A cached number is a Tuesday audit

,

The same failure mode applies to statistics. An answer engine recently summarised this site and quoted our public scan counter — accurately, for the day it had read the page, which put it about 140 scans behind the live figure by the time the summary was shown. Nothing lied; a snapshot aged. It is the whole argument of this post applied to a single number, and it is why every figure on this page is recomputed from the live record at load time and why the citable unit here is the URL, not the quote.

,

What this sample is, stated plainly

,

The panel is seeded and submitted — sites we chose to watch and sites people asked us to scan — not a random sample of the web. Quote it as “among the sites CrawlCheck measures daily”, nothing broader. No scanned domain is named here or anywhere else on this site.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

All findings · The dataset · How the dataset works