CrawlCheck

Library · Scanner corrections

Scanner corrections

In one sentence

Findings are short case studies written from real measurements — a robots.txt that answered 200 with a challenge page, a site whose GPTBot traffic came from one address wearing seven crawler names, a homepage that was two percent readable text, a contact email that bounced for one published character — each with the bytes that showed it, and none naming a scanned third-party domain.

Where our own scanner, report or pipeline was wrong, and the fix.

71findings
34guides
47finding codes explained
2,474scans behind the live figures

2026-09-30 · Scanner corrections4 min

Five days of our alerts went to an address with no mailbox

Every alert was accepted by the sending API and dropped by a suppression list, because the operator address had no mailbox after a mail migration. What passed, what was measured, and the canary that would have caught it.

Read the measurements →

2026-09-30 · Scanner corrections4 min

Our security headers skipped every signed-in session

A CSP and a CORS tightening were verified live and called done. For every signed-in session neither was applied, because one async call was not awaited. The defect, the miss, and the rule.

Read the measurements →

2026-09-30 · Scanner corrections5 min

Our extension read a page as 36.5% JavaScript. It was zero.

How a comparison of two copies of the same 2,850 words came out a third wrong: undecoded entities and a signed-in admin bar. What version 1.0.2 compares instead, and the controls that prove it.

Read the measurements →

2026-09-30 · Scanner corrections5 min

Eight contradictions in our own copy, rechecked today

An outside audit listed eight self-contradictions on crawlcheck.io. Each one rechecked against the live site: what the copy says now, which were real, and the one still wrong this morning.

Read the measurements →

2026-09-25 · Scanner corrections5 min

We failed our own contrast check. Hover states were worse.

One invisible line in a screenshot led to 55,118 measurements at rest, then 1,878 hovers and 1,834 keyboard focuses. What failed, why static checks missed the worst of it, and what we changed.

Read the measurements →

2026-09-23 · Scanner corrections6 min

Authorization bug: why our owner's scans came back locked

The scan endpoint knew who he was and used it to skip the rate limit, then gated the reply as if he were a stranger. Two gates asked the same question two different ways.

Read the measurements →

2026-09-17 · Scanner corrections6 min

Is your host blocking ClaudeBot? Group the 403s by IP first

A 57% block rate in our own access logs looked like a hosting provider shutting out AI crawlers. Grouped by source address, almost all of it was one machine wearing seven different crawler names.

Read the measurements →

2026-09-17 · Scanner corrections5 min

When an LLM summarises your audit report and gets the grade wrong

It decided we were docking a site's grade for things we admitted we could not measure. We were not. But our report made that the natural reading, and that is our defect, not the reader's.

Read the measurements →

2026-09-17 · Scanner corrections5 min

Our AI-crawler forgery rate was counting our own browser extension

Our extension forges a crawler user-agent on every probe — by design. To this site's own telemetry that is indistinguishable from a stranger doing the same thing, so for weeks our users inflated the number we publish about everyone else.

Read the measurements →

2026-09-17 · Scanner corrections5 min

We put the findings behind a paywall and the report started saying there were none

Emptying an array hid the fix list everywhere at once — and made the page and the CSV assert that a site with findings had none. What an empty collection means when it can mean two things.

Read the measurements →

2026-09-17 · Scanner corrections5 min

Our control panel answered 200 with nothing in it

A function called a block with an argument it never received. The route caught the error, replaced the panel with one sentence, and kept returning 200 — so every check we had passed.

Read the measurements →

2026-09-16 · Scanner corrections5 min

Our crawl-waste check named a parameter called amp;plan. No site has one.

Nineteen parameter names in the corpus began with amp; because an href was split on & before its escaped ampersands were decoded. The counts were right the whole time. The names beside them were not.

Read the measurements →

2026-09-12 · Scanner corrections5 min

Three of our own rows passed because we measured nothing

A scan that read no file at all found zero disagreements between hosts, passed the only scored row in that section, and could report 100 out of 100 for a site we never managed to read.

Read the measurements →

2026-09-09 · Scanner corrections6 min

Our scanner counted 14 placeholders and none of the 13 photographs

A lazyload plugin puts a grey placeholder in src and the real URL one attribute over. Our image reader took the placeholder at face value and reported a page of webp photographs as having no modern images at all.

Read the measurements →

2026-09-08 · Scanner corrections5 min

Our proofs said “pending” for 26 days. The anchoring was fine.

Every timestamp file this site served was a calendar receipt rather than a finished Bitcoin proof. The record was unbroken the whole time; the artifact handed to a reader did not show it.

Read the measurements →

2026-09-08 · Scanner corrections4 min

Our scanner said the viewport tag was missing. It was there all along.

A page scored 0 out of 100 on the mobile section with a valid viewport meta in its head. The attribute was unquoted, and four of our extractors only matched quoted attributes.

Read the measurements →

2026-09-01 · Scanner corrections5 min

Firewall or website? When an AI visibility scanner grades your WAF

We published a failing grade for a competitor. The grade was wrong, and the way it was wrong is the most useful thing this scanner has taught us.

Read the measurements →

2026-09-01 · Scanner corrections5 min

Why a self-identifying crawler header must never clear a check

A marker anyone can send must never be able to clear a check. It can label a request; it can never absolve one.

Read the measurements →

2026-09-01 · Scanner corrections3 min

A Worker cannot fetch its own host: the sameAs profile we called dead

A customer's schema pointed at a page on our own domain. From outside it answered 200. From inside the scanner it answered nothing, and we reported it dead.

Read the measurements →

2026-09-01 · Scanner corrections2 min

Our placeholder check accused a newspaper of publishing NaN. The NaN was in a URL.

A self-audit rule that scans every stored record for the token NaN found three of them. All three were inside the letters of one URL slug.

Read the measurements →

2026-09-01 · Scanner corrections4 min

We reviewed all 150 scored rows against seven site types. Here is what we unscored, and why.

A threshold that fails a lean personal blog is not measuring quality. A row that scores the mechanism when another row scores the outcome counts one defect twice.

Read the measurements →

2026-08-26 · Scanner corrections5 min

Our audit labelled one page and measured another for weeks

Every scan of a specific URL actually measured the homepage. The report header showed the right path. The numbers underneath belonged to a different page.

Read the measurements →

2026-08-21 · Scanner corrections3 min

CRLF corruption: a save that returned success and rewrote every newline

The write reported success. Read back, the file was 16,330 bytes larger — exactly the number of newlines it contained. A store confirming it accepted your bytes is not the same as confirming it kept them.

Read the measurements →

2026-08-21 · Scanner corrections4 min

An analytics script added 650 ms to every request: how we found it

A first-party beacon reported that 70% of our LCP was time-to-first-byte. The cause was three telemetry writes awaited before routing — next to a comment explaining why they could not be moved. The comment was wrong.

Read the measurements →

2026-08-20 · Scanner corrections4 min

Cache negative lookups too: a miss that is never cached repays forever

A cache that only records successes speeds up the requests that were already fast and does nothing for the ones that hurt.

Read the measurements →

2026-08-20 · Scanner corrections3 min

Duplicate @id in JSON-LD: one node parsed twice, not two businesses

An @id is the identifier. Two records carrying the same one are, by definition, the same record appearing twice.

Read the measurements →

Questions about this page

QAre these real sites?
Yes, every finding is a measurement taken on a live site. Domains are not named unless the author owns them.
QHow is a finding different from a guide?
A finding reports one measured case and what it showed. A guide explains how to run and read the check yourself.

Keep reading