Library · Scanner corrections
Scanner corrections
In one sentence
Findings are short case studies written from real measurements — a robots.txt that answered 200 with a challenge page, a site whose GPTBot traffic came from one address wearing seven crawler names, a homepage that was two percent readable text, a contact email that bounced for one published character — each with the bytes that showed it, and none naming a scanned third-party domain.
Where our own scanner, report or pipeline was wrong, and the fix.
2026-09-30 · Scanner corrections4 min
Five days of our alerts went to an address with no mailbox
Every alert was accepted by the sending API and dropped by a suppression list, because the operator address had no mailbox after a mail migration. What passed, what was measured, and the canary that would have caught it.
Read the measurements →
2026-09-30 · Scanner corrections4 min
Our security headers skipped every signed-in session
A CSP and a CORS tightening were verified live and called done. For every signed-in session neither was applied, because one async call was not awaited. The defect, the miss, and the rule.
Read the measurements →
2026-09-30 · Scanner corrections5 min
Our extension read a page as 36.5% JavaScript. It was zero.
How a comparison of two copies of the same 2,850 words came out a third wrong: undecoded entities and a signed-in admin bar. What version 1.0.2 compares instead, and the controls that prove it.
Read the measurements →
2026-09-30 · Scanner corrections5 min
Eight contradictions in our own copy, rechecked today
An outside audit listed eight self-contradictions on crawlcheck.io. Each one rechecked against the live site: what the copy says now, which were real, and the one still wrong this morning.
Read the measurements →
2026-09-25 · Scanner corrections5 min
We failed our own contrast check. Hover states were worse.
One invisible line in a screenshot led to 55,118 measurements at rest, then 1,878 hovers and 1,834 keyboard focuses. What failed, why static checks missed the worst of it, and what we changed.
Read the measurements →
2026-09-23 · Scanner corrections6 min
Authorization bug: why our owner's scans came back locked
The scan endpoint knew who he was and used it to skip the rate limit, then gated the reply as if he were a stranger. Two gates asked the same question two different ways.
Read the measurements →
2026-09-17 · Scanner corrections6 min
Is your host blocking ClaudeBot? Group the 403s by IP first
A 57% block rate in our own access logs looked like a hosting provider shutting out AI crawlers. Grouped by source address, almost all of it was one machine wearing seven different crawler names.
Read the measurements →
2026-09-17 · Scanner corrections5 min
When an LLM summarises your audit report and gets the grade wrong
It decided we were docking a site's grade for things we admitted we could not measure. We were not. But our report made that the natural reading, and that is our defect, not the reader's.
Read the measurements →
2026-09-17 · Scanner corrections5 min
Our AI-crawler forgery rate was counting our own browser extension
Our extension forges a crawler user-agent on every probe — by design. To this site's own telemetry that is indistinguishable from a stranger doing the same thing, so for weeks our users inflated the number we publish about everyone else.
Read the measurements →
2026-09-17 · Scanner corrections5 min
We put the findings behind a paywall and the report started saying there were none
Emptying an array hid the fix list everywhere at once — and made the page and the CSV assert that a site with findings had none. What an empty collection means when it can mean two things.
Read the measurements →
2026-09-17 · Scanner corrections5 min
Our control panel answered 200 with nothing in it
A function called a block with an argument it never received. The route caught the error, replaced the panel with one sentence, and kept returning 200 — so every check we had passed.
Read the measurements →
2026-09-16 · Scanner corrections5 min
Our crawl-waste check named a parameter called amp;plan. No site has one.
Nineteen parameter names in the corpus began with amp; because an href was split on & before its escaped ampersands were decoded. The counts were right the whole time. The names beside them were not.
Read the measurements →
2026-09-12 · Scanner corrections5 min
Three of our own rows passed because we measured nothing
A scan that read no file at all found zero disagreements between hosts, passed the only scored row in that section, and could report 100 out of 100 for a site we never managed to read.
Read the measurements →
2026-09-09 · Scanner corrections6 min
Our scanner counted 14 placeholders and none of the 13 photographs
A lazyload plugin puts a grey placeholder in src and the real URL one attribute over. Our image reader took the placeholder at face value and reported a page of webp photographs as having no modern images at all.
Read the measurements →
2026-09-08 · Scanner corrections5 min
Our proofs said “pending” for 26 days. The anchoring was fine.
Every timestamp file this site served was a calendar receipt rather than a finished Bitcoin proof. The record was unbroken the whole time; the artifact handed to a reader did not show it.
Read the measurements →
2026-09-08 · Scanner corrections4 min
Our scanner said the viewport tag was missing. It was there all along.
A page scored 0 out of 100 on the mobile section with a valid viewport meta in its head. The attribute was unquoted, and four of our extractors only matched quoted attributes.
Read the measurements →
2026-09-01 · Scanner corrections5 min
Firewall or website? When an AI visibility scanner grades your WAF
We published a failing grade for a competitor. The grade was wrong, and the way it was wrong is the most useful thing this scanner has taught us.
Read the measurements →
2026-09-01 · Scanner corrections5 min
Why a self-identifying crawler header must never clear a check
A marker anyone can send must never be able to clear a check. It can label a request; it can never absolve one.
Read the measurements →
2026-09-01 · Scanner corrections3 min
A Worker cannot fetch its own host: the sameAs profile we called dead
A customer's schema pointed at a page on our own domain. From outside it answered 200. From inside the scanner it answered nothing, and we reported it dead.
Read the measurements →
2026-09-01 · Scanner corrections2 min
Our placeholder check accused a newspaper of publishing NaN. The NaN was in a URL.
A self-audit rule that scans every stored record for the token NaN found three of them. All three were inside the letters of one URL slug.
Read the measurements →
2026-09-01 · Scanner corrections4 min
We reviewed all 150 scored rows against seven site types. Here is what we unscored, and why.
A threshold that fails a lean personal blog is not measuring quality. A row that scores the mechanism when another row scores the outcome counts one defect twice.
Read the measurements →
2026-08-26 · Scanner corrections5 min
Our audit labelled one page and measured another for weeks
Every scan of a specific URL actually measured the homepage. The report header showed the right path. The numbers underneath belonged to a different page.
Read the measurements →
2026-08-21 · Scanner corrections3 min
CRLF corruption: a save that returned success and rewrote every newline
The write reported success. Read back, the file was 16,330 bytes larger — exactly the number of newlines it contained. A store confirming it accepted your bytes is not the same as confirming it kept them.
Read the measurements →
2026-08-21 · Scanner corrections4 min
An analytics script added 650 ms to every request: how we found it
A first-party beacon reported that 70% of our LCP was time-to-first-byte. The cause was three telemetry writes awaited before routing — next to a comment explaining why they could not be moved. The comment was wrong.
Read the measurements →
2026-08-20 · Scanner corrections4 min
Cache negative lookups too: a miss that is never cached repays forever
A cache that only records successes speeds up the requests that were already fast and does nothing for the ones that hurt.
Read the measurements →
2026-08-20 · Scanner corrections3 min
Duplicate @id in JSON-LD: one node parsed twice, not two businesses
An @id is the identifier. Two records carrying the same one are, by definition, the same record appearing twice.
Read the measurements →