CrawlCheck

Findings · 2026-08-18 · By · 0 views

When an LLM summarises your audit report and gets the grade wrong

It decided we were docking a site's grade for things we admitted we could not measure. We were not. But our report made that the natural reading, and that is our defect, not the reader's.

A score that cannot name the input that produced it is not a measurement. Our grade capped a site at C without stating which finding did it, so the letter was unarguable — and an unarguable number is indistinguishable from an arbitrary one. Every stored record now carries the finding that set the cap, so the grade can be checked rather than trusted.

Someone ran one of our reports through a language model for a second opinion. It found a single flaw: the site scored 90 out of 100, sat at the 86th percentile of our corpus, and still graded C. It concluded the letter was being docked for categories the tool itself admitted it could not measure — Core Web Vitals with no field data, and agent surfaces still too new to score.

That would be a genuinely bad design. Penalising an operator for your measurement gaps punishes them for your limitations.

It is not what happened #

The letter is capped by the most severe finding on the site, and only by findings. That scan carried exactly one finding, of medium severity: the homepage was served from an edge cache copy nearly four hours old. That finding capped 90 at C. The unmeasured categories never touched it.

Why the reader was right anyway #

Our report placed the cap and the list of unscored sections side by side and said nothing about which caused which. Adjacency implied causation. A careful reader followed the implication to a conclusion the data did not support, because the data never ruled it out.

For a product whose entire pitch is legible evidence, that is a defect in the output, not in the reader. Reports now state the reason for a cap — which finding set it — and state that unscored sections did not affect the grade.

Three states, and only one of them is a zero #

The confusion had a root deeper than the field order. A section can end in three states, and collapsing any two of them produces exactly this kind of misreading:

StateWhat it meansEffect on the score
Measured, passedThe check ran and the site met itFull weight
Measured, failedThe check ran and the site did not meet itScored down, and a finding if severe enough
UnmeasuredNo field data, no applicable surface, or the input never arrivedExcluded from the denominator. Not a zero.

An unmeasured component scored as zero is the single most common way an audit tool manufactures a bad grade. It is why a brand-new site with no field performance data can appear to be failing at performance, and why a service-area business with no street address can appear to be hiding one. Here the row is dropped and the report names it — the same three-outcome rule that governs every unverifiable claim.

The published count now reads the scorer #

A related drift, found later and worth naming because it is the same disease: the site’s own marketing copy said one number of scored sections while the scorer used another. The count was hand-typed, so the two could disagree indefinitely and nothing would notice. Today the published figure is computed from the scorer itself, so it cannot drift from the thing it describes. As you read this it reports 35 sections, 24 of them scored; the other 11 are measured and shown without touching the grade — read live as this page loads, not typed here.

That is the same fix as stating the cap’s reason, applied one level up. A number that describes another number has to be computed from it, not maintained beside it.

The same defect, found twice more #

Once you go looking for numbers that depend on other numbers without saying so, they turn up everywhere. Two more from this product, both fixed the same way:

Both are the same shape as the grade cap: a reader, human or machine, cannot tell from the output which fields are load-bearing. The cure is not more prose. It is putting the relationship in the data, where it can be checked by whoever is reading — including the published dataset, which states its method rather than describing it in a paragraph elsewhere.

The general lesson #

Increasingly your reports are read by machines as well as people, and a machine has no access to the paragraph of context you would have supplied in a meeting. Every number that changes another number has to say so in the data, not in the prose beside it. A field that needs context to be read correctly will eventually be read without it — and the reading will be confident.

The test is cheap to run on your own output: hand a report to a model, ask it to explain the headline number, and see whether the explanation is right. If it invents a plausible cause, your data implied one. That is not the model hallucinating. That is your schema.

The fields that now travel with every grade #

Since this incident, every stored record carries cap_reason naming the finding that capped the letter, unscored_affects_grade: false stated rather than implied, the list of sections that were excluded and why, and the score version the numbers were computed under. A refused scan carries no grade at all rather than a low one, after the site that scored 15 because it refused us. Anything that changes a number says so next to the number.

How to check whether your own reports explain themselves #

Take any report your organisation produces, remove every sentence of prose, and hand the remaining fields to someone who has never seen it. If they can reconstruct why the headline number is what it is, the data is legible. If they need the paragraph beside it, a language model summarising that report will get it wrong in the same place a careful human did here. The glossary entry for evidence level describes the vocabulary this site uses to say how strongly a thing is known; how to measure AI visibility covers what a grade here is a grade of.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Can a report be accurate and still mislead?

Yes. Every number can be correct while the layout implies a relationship that does not exist. A reader who concludes the wrong thing from a truthful page has been misled by the presentation.

What did the reader get wrong, and why was it our fault?

It read unscored rows as deductions. The rows said they were unscored, but nothing in the design separated them from the ones that counted.

What is the general lesson for report design?

If a caveat only survives when someone reads carefully, it is decoration. Structure has to carry it.

What should a score always carry?

The input that produced it. If a letter is capped, the finding that capped it; if a section is unscored, the reason. A number that cannot be traced back to an observation cannot be disputed, and a number nobody can dispute is not evidence.

Does an unscored section lower a grade?

No. Sections measured but not yet scored are stated as unscored and excluded from the arithmetic, and the record says so explicitly rather than leaving a reader to infer it.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works