CrawlCheck

Findings · 2026-08-15 · By · 0 views

Generic AEO scores penalise sites for checks that do not apply to them

A public checker deducted seven points from a tree service for having no OAuth server, no payment endpoint and no MCP card — after its own output had already said the site was not a commerce site.

A checker that gives every site the same list is measuring the list, not the site. That is the whole disagreement, and it is worth setting out plainly rather than as a complaint about anybody’s product.

The case that made this concrete #

A public agent-readiness checker was pointed at a Denver tree service. It returned 29 out of 100. Seven of the missing points sat in one category: no OAuth authorization server, no x402 payment endpoint, no MPP OpenAPI document, no UCP profile, no ACP document, no MCP card, no API catalog.

The interesting part is not the score. It is that the same output said “not a commerce site” next to four of those items — and counted them anyway. The tool had already worked out that the checks did not apply, and then deducted for them regardless.

Scanned here, the same site grades A. Not because this scanner is kinder. Because a tree service having no payment-protocol endpoint is not a defect, and a measurement that treats it as one is answering a question nobody asked.

What we do instead, stated so you can check it #

Every emerging-standard surface is measured and none of them are scored. They appear in the report as present or absent with the evidence attached, and they move no grade in either direction. When a surface does not apply, the row says why in words rather than silently subtracting.

How exclusion works, mechanically #

Three outcomes exist for every scored row and they are never collapsed: true, false, and null. A null row is one the site gave no evidence for either way, and it is left out of the section’s arithmetic entirely rather than counted as a failure. A section whose rows are all null is not scored at all; the report shows the evidence and a stated reason, and the grade is computed from the sections that remain. That is the rule the checker above broke: it had the reason (“not a commerce site”) and scored the row false anyway.

SectionScored whenOtherwise
commerceProduct or Offer schema, or a checkout, is on the pageexcluded; the row names what was not found
localthe page is about a local business — a LocalBusiness node whose own URL is this site, or exactly one declaredlocal:false with the reason and the nodes that were declared
vitalsGoogle has field data for the originexcluded; not simulated here
entity corroborationat least one sameAs target could be fetched and judgedexcluded; unverifiable is not unreciprocated
freshness, crawl waste, entity paritynever — measured, published, weight 0reported with evidence; a candidate for scoring, not a score

Every weight is declared in one map; the published “sections scored” count is derived from it, so the number on the homepage cannot drift from the scorer.

The time this scanner got the same thing wrong #

The rule is easy to state and it was still broken here, so the case belongs in this post. The local section used to take the first LocalBusiness node in a page’s graph as the subject. A personal site that lists the nine trade brands its owner runs declares nine such nodes, none of which is the site itself. The scanner graded that page as whichever company came first in the markup and marked it down for that company having no address, hours or phone on a page that was never about it. Local read 17 out of 100. The fix was the same principle as the commerce rule: a page that describes local businesses is not one. The section is now excluded on that site with the reason stated, and the grade moved from A/91 to A/94 with no change to the page. The three real local-business sites used as controls did not move at all.

The general form of that defect is worth naming. Every scored row was read against seven sites for exactly this reason, and several rows were unscored because they described the corpus rather than the site in front of the scanner. A grade that cannot explain itself in terms of the site’s own evidence is the same failure from the other side.

Applicability, measured live #

The table below is not an opinion about adoption. It is read from this scanner’s own dataset when the page loads, across 5752 scans, and it changes as the corpus grows.

SurfaceSites serving itShare of 5752 scans
agents.md
instructions written for an agent rather than a crawler
159327.7%
Link headers
registered relations only — describedby, service-desc
284149.4%
Markdown negotiation
a plain-text form of the page for a client that asks for one
117720.5%

Read from this scanner’s own counters when the page loaded. None of these three is scored here, in either direction — they are reported with their evidence and left alone.

Two of those rows are low, and that is the point of publishing them. A surface that almost nobody serves is not a stick to beat a small business with — it is a description of where the standard actually is. A scoring model that treats a 5% adoption rate as a 95% failure rate is measuring the calendar, not the website.

Where this scanner is weaker, since the comparison should run both ways #

A single scan here reports what is true now. Page-level sections describe the home page; a scan requested for a deep path says so on the report and is deliberately not recorded, because filing a page under the site would corrupt a history that cannot be rebuilt. Core Web Vitals come from Google’s public field data, not from a test run here, and are absent when Google has no sample for the site. None of that is hidden behind a paid tier.

The claim, stated narrowly #

The claim being made is narrow and we would rather it be checkable than impressive: applicability is derived from the site rather than assumed, and nothing a payment unlocks changes a measurement, a finding or a grade. The weights, the section list and the scored count are all readable from the public endpoints, which is how you check it.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Why do generic AEO scores penalise small sites?

Because they apply every check to every site. A tree service was docked seven points for having no OAuth server, no payment endpoint and no MCP card.

What is the alternative to a universal checklist?

Deriving applicability from the site itself. A check that does not apply should say so and score nothing, rather than scoring zero.

Where is this scanner weaker than the tool it criticises?

It runs fewer rendering checks, and it says so in the report rather than leaving the gap for a reader to discover.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works