Findings · 2026-08-15 · By VSNARY | Emmanuel Orta · 0 views
Generic AEO scores penalise sites for checks that do not apply to them
A public checker deducted seven points from a tree service for having no OAuth server, no payment endpoint and no MCP card — after its own output had already said the site was not a commerce site.
A checker that gives every site the same list is measuring the list, not the site. That is the whole disagreement, and it is worth setting out plainly rather than as a complaint about anybody’s product.
The case that made this concrete #
A public agent-readiness checker was pointed at a Denver tree service. It returned 29 out of 100. Seven of the missing points sat in one category: no OAuth authorization server, no x402 payment endpoint, no MPP OpenAPI document, no UCP profile, no ACP document, no MCP card, no API catalog.
The interesting part is not the score. It is that the same output said “not a commerce site” next to four of those items — and counted them anyway. The tool had already worked out that the checks did not apply, and then deducted for them regardless.
Scanned here, the same site grades A. Not because this scanner is kinder. Because a tree service having no payment-protocol endpoint is not a defect, and a measurement that treats it as one is answering a question nobody asked.
What we do instead, stated so you can check it #
Every emerging-standard surface is measured and none of them are scored. They appear in the report as present or absent with the evidence attached, and they move no grade in either direction. When a surface does not apply, the row says why in words rather than silently subtracting.
- Payment protocols are evaluated only where the site has Product or Offer schema, or a detected checkout. Otherwise the row reads: no Product, Offer or checkout found.
- OAuth is evaluated only where an API catalog was found. An authorization server for an API that does not exist is not a gap.
- Web Bot Auth is described from the caller side, because it identifies a site that sends signed agent requests, not one that receives traffic.
- Link headers are checked against registered IANA relations only. Inventing a relation to satisfy a checker would produce a better score and a worse site.
How exclusion works, mechanically #
Three outcomes exist for every scored row and they are never collapsed: true, false, and null. A null row is one the site gave no evidence for either way, and it is left out of the section’s arithmetic entirely rather than counted as a failure. A section whose rows are all null is not scored at all; the report shows the evidence and a stated reason, and the grade is computed from the sections that remain. That is the rule the checker above broke: it had the reason (“not a commerce site”) and scored the row false anyway.
| Section | Scored when | Otherwise |
|---|---|---|
| commerce | Product or Offer schema, or a checkout, is on the page | excluded; the row names what was not found |
| local | the page is about a local business — a LocalBusiness node whose own URL is this site, or exactly one declared | local:false with the reason and the nodes that were declared |
| vitals | Google has field data for the origin | excluded; not simulated here |
| entity corroboration | at least one sameAs target could be fetched and judged | excluded; unverifiable is not unreciprocated |
| freshness, crawl waste, entity parity | never — measured, published, weight 0 | reported with evidence; a candidate for scoring, not a score |
Every weight is declared in one map; the published “sections scored” count is derived from it, so the number on the homepage cannot drift from the scorer.
The time this scanner got the same thing wrong #
The rule is easy to state and it was still broken here, so the case belongs in this post. The local section used to take the first LocalBusiness node in a page’s graph as the subject. A personal site that lists the nine trade brands its owner runs declares nine such nodes, none of which is the site itself. The scanner graded that page as whichever company came first in the markup and marked it down for that company having no address, hours or phone on a page that was never about it. Local read 17 out of 100. The fix was the same principle as the commerce rule: a page that describes local businesses is not one. The section is now excluded on that site with the reason stated, and the grade moved from A/91 to A/94 with no change to the page. The three real local-business sites used as controls did not move at all.
The general form of that defect is worth naming. Every scored row was read against seven sites for exactly this reason, and several rows were unscored because they described the corpus rather than the site in front of the scanner. A grade that cannot explain itself in terms of the site’s own evidence is the same failure from the other side.
Applicability, measured live #
The table below is not an opinion about adoption. It is read from this scanner’s own dataset when the page loads, across 5752 scans, and it changes as the corpus grows.
| Surface | Sites serving it | Share of 5752 scans |
|---|---|---|
| agents.md instructions written for an agent rather than a crawler | 1593 | 27.7% |
| Link headers registered relations only — describedby, service-desc | 2841 | 49.4% |
| Markdown negotiation a plain-text form of the page for a client that asks for one | 1177 | 20.5% |
Read from this scanner’s own counters when the page loaded. None of these three is scored here, in either direction — they are reported with their evidence and left alone.
Two of those rows are low, and that is the point of publishing them. A surface that almost nobody serves is not a stick to beat a small business with — it is a description of where the standard actually is. A scoring model that treats a 5% adoption rate as a 95% failure rate is measuring the calendar, not the website.
Where this scanner is weaker, since the comparison should run both ways #
A single scan here reports what is true now. Page-level sections describe the home page; a scan requested for a deep path says so on the report and is deliberately not recorded, because filing a page under the site would corrupt a history that cannot be rebuilt. Core Web Vitals come from Google’s public field data, not from a test run here, and are absent when Google has no sample for the site. None of that is hidden behind a paid tier.
The claim, stated narrowly #
The claim being made is narrow and we would rather it be checkable than impressive: applicability is derived from the site rather than assumed, and nothing a payment unlocks changes a measurement, a finding or a grade. The weights, the section list and the scored count are all readable from the public endpoints, which is how you check it.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
Why do generic AEO scores penalise small sites?
Because they apply every check to every site. A tree service was docked seven points for having no OAuth server, no payment endpoint and no MCP card.
What is the alternative to a universal checklist?
Deriving applicability from the site itself. A check that does not apply should say so and score nothing, rather than scoring zero.
Where is this scanner weaker than the tool it criticises?
It runs fewer rendering checks, and it says so in the report rather than leaving the gap for a reader to discover.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.