Findings · 2026-08-18 · By VSNARY | Emmanuel Orta · 0 views
When an LLM summarises your audit report and gets the grade wrong
It decided we were docking a site's grade for things we admitted we could not measure. We were not. But our report made that the natural reading, and that is our defect, not the reader's.
A score that cannot name the input that produced it is not a measurement. Our grade capped a site at C without stating which finding did it, so the letter was unarguable — and an unarguable number is indistinguishable from an arbitrary one. Every stored record now carries the finding that set the cap, so the grade can be checked rather than trusted.
Someone ran one of our reports through a language model for a second opinion. It found a single flaw: the site scored 90 out of 100, sat at the 86th percentile of our corpus, and still graded C. It concluded the letter was being docked for categories the tool itself admitted it could not measure — Core Web Vitals with no field data, and agent surfaces still too new to score.
That would be a genuinely bad design. Penalising an operator for your measurement gaps punishes them for your limitations.
It is not what happened #
The letter is capped by the most severe finding on the site, and only by findings. That scan carried exactly one finding, of medium severity: the homepage was served from an edge cache copy nearly four hours old. That finding capped 90 at C. The unmeasured categories never touched it.
Why the reader was right anyway #
Our report placed the cap and the list of unscored sections side by side and said nothing about which caused which. Adjacency implied causation. A careful reader followed the implication to a conclusion the data did not support, because the data never ruled it out.
For a product whose entire pitch is legible evidence, that is a defect in the output, not in the reader. Reports now state the reason for a cap — which finding set it — and state that unscored sections did not affect the grade.
Three states, and only one of them is a zero #
The confusion had a root deeper than the field order. A section can end in three states, and collapsing any two of them produces exactly this kind of misreading:
| State | What it means | Effect on the score |
|---|---|---|
| Measured, passed | The check ran and the site met it | Full weight |
| Measured, failed | The check ran and the site did not meet it | Scored down, and a finding if severe enough |
| Unmeasured | No field data, no applicable surface, or the input never arrived | Excluded from the denominator. Not a zero. |
An unmeasured component scored as zero is the single most common way an audit tool manufactures a bad grade. It is why a brand-new site with no field performance data can appear to be failing at performance, and why a service-area business with no street address can appear to be hiding one. Here the row is dropped and the report names it — the same three-outcome rule that governs every unverifiable claim.
The published count now reads the scorer #
A related drift, found later and worth naming because it is the same disease: the site’s own marketing copy said one number of scored sections while the scorer used another. The count was hand-typed, so the two could disagree indefinitely and nothing would notice. Today the published figure is computed from the scorer itself, so it cannot drift from the thing it describes. As you read this it reports 35 sections, 24 of them scored; the other 11 are measured and shown without touching the grade — read live as this page loads, not typed here.
That is the same fix as stating the cap’s reason, applied one level up. A number that describes another number has to be computed from it, not maintained beside it.
The same defect, found twice more #
Once you go looking for numbers that depend on other numbers without saying so, they turn up everywhere. Two more from this product, both fixed the same way:
- A section reported as measured with nothing behind it. A check whose input never arrived looked the same as one that ran and found nothing. Ran, found none and did not run, here is why are different sentences, and the report now shows them differently.
- A display field read as a scored one. Two image measurements were wrong for months and no grade moved, because they were reported rather than scored. A wrong number nobody loses points for is harder to catch, not less wrong.
Both are the same shape as the grade cap: a reader, human or machine, cannot tell from the output which fields are load-bearing. The cure is not more prose. It is putting the relationship in the data, where it can be checked by whoever is reading — including the published dataset, which states its method rather than describing it in a paragraph elsewhere.
The general lesson #
Increasingly your reports are read by machines as well as people, and a machine has no access to the paragraph of context you would have supplied in a meeting. Every number that changes another number has to say so in the data, not in the prose beside it. A field that needs context to be read correctly will eventually be read without it — and the reading will be confident.
The test is cheap to run on your own output: hand a report to a model, ask it to explain the headline number, and see whether the explanation is right. If it invents a plausible cause, your data implied one. That is not the model hallucinating. That is your schema.
The fields that now travel with every grade #
Since this incident, every stored record carries cap_reason naming the finding that capped the letter, unscored_affects_grade: false stated rather than implied, the list of sections that were excluded and why, and the score version the numbers were computed under. A refused scan carries no grade at all rather than a low one, after the site that scored 15 because it refused us. Anything that changes a number says so next to the number.
How to check whether your own reports explain themselves #
Take any report your organisation produces, remove every sentence of prose, and hand the remaining fields to someone who has never seen it. If they can reconstruct why the headline number is what it is, the data is legible. If they need the paragraph beside it, a language model summarising that report will get it wrong in the same place a careful human did here. The glossary entry for evidence level describes the vocabulary this site uses to say how strongly a thing is known; how to measure AI visibility covers what a grade here is a grade of.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
Can a report be accurate and still mislead?
Yes. Every number can be correct while the layout implies a relationship that does not exist. A reader who concludes the wrong thing from a truthful page has been misled by the presentation.
What did the reader get wrong, and why was it our fault?
It read unscored rows as deductions. The rows said they were unscored, but nothing in the design separated them from the ones that counted.
What is the general lesson for report design?
If a caveat only survives when someone reads carefully, it is decoration. Structure has to carry it.
What should a score always carry?
The input that produced it. If a letter is capped, the finding that capped it; if a section is unscored, the reason. A number that cannot be traced back to an observation cannot be disputed, and a number nobody can dispute is not evidence.
Does an unscored section lower a grade?
No. Sections measured but not yet scored are stated as unscored and excluded from the arithmetic, and the record says so explicitly rather than leaving a reader to infer it.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.