Methodology
The rulebook
Every finding the scanner can raise: what it means, how serious it is, how to fix it and how many scans carried it. Open a rule for the detail.
/api/rules (JSON) · schema · compatibility policy · trust
Each finding has a level from info to critical. Open findings at medium and above lower the letter grade; fix them first. A rule that changes what it reports gets a new revision; a record keeps the revision that decided it, and a replay uses that revision, never today's.
Every rule
Site, hosts and page 8 Headers 4 Mostly code 3 Machine files 11 Robots 13 Delivery parity 8 JSON-LD and entities 3 NAP and claims 2 Answer-engine citations 5 Capabilities 6 Self-audit: defects in our own record 12 MCP servers 1 High-risk field watch 1
Site, hosts and page
CSP_REPORTING_BROKENCSP violation reporting is misconfiguredinfo0.4%rev 1
- Why it matters
- The Content-Security-Policy header declares where to send violation reports, and the declaration is wrong. `report-to` takes the NAME of a group defined in a separate header — given a URL it is silently ignored, so nothing is reported at all. And `report-uri` aimed at your own origin turns every violation into another request to the page you are already trying to keep fast. Neither breaks the page; both mean the reporting you think you have is either absent or self-inflicted.
- Fix
- The Content-Security-Policy report endpoint does not accept reports. Fix the report-uri / report-to target or remove the directive.
effort: config - Share of scans
- 0.4% — 11 of 3137 counted scans
- Links
- JSON · glossary: content-security-policy
BROKEN_INTERNAL_LINKLinks on this site point at pages that do not answermedium3%rev 1
- Why it matters
- A link on this site was followed and the target answered 4xx or 5xx. A crawler spends budget on every one of these and learns nothing; a reader following one gets an error page. The usual causes are a route renamed without updating the links to it, and a nav or footer that renders links to authenticated-only routes for everyone — those answer 404 when signed out, which is what a crawler always is. Only 404 and 410 count here. A link answering 401, 403 or 429 means something declined to answer this scanner — bot protection, a login wall, a rate limit — which is a fact about the request, not about the link, so those are counted and never reported as broken. Links are sampled rather than exhaustively crawled, so this is a floor and not a full link audit.
- Fix
- Fix or remove the internal links that return 404 so crawlers and visitors are not sent to dead pages.
effort: content - Share of scans
- 3% — 94 of 3137 counted scans
- Links
- JSON · glossary: broken-link · glossary: internal-link · glossary: dead-anchor
HOST_DUPLICATE_200www and the apex both serve the site instead of one redirectingmedium3.5%rev 1
- Why it matters
- Both www and the bare domain answer 200 with the same pages, so every URL on this site exists at two addresses and a crawler fetches the whole site twice. A canonical tag does not prevent this: it is read after the response arrives, so the second fetch has already been spent. Nothing looks broken from a browser, which is why this survives for years. The fix is a 301 from one host to the other so there is one address per page.
- Fix
- Pick one host (apex or www) and redirect the other to it with a 301 so there is one copy of every page.
effort: config - Share of scans
- 3.5% — 111 of 3137 counted scans
- Links
- JSON · glossary: canonical · glossary: canonical-url · glossary: host-duplicate
ORIGIN_REFUSED_SCANNERThe homepage and robots.txt both refused this client, so the site was never measuredcritical0.3%rev 1
- Fix
- Allow this scanner's address through the security layer, or scan from the browser extension, which uses your own address.
effort: config - Share of scans
- 0.3% — 8 of 3137 counted scans
- Links
- JSON · glossary: refusal · glossary: datacenter-ip · glossary: vantage-point
DOMAIN_IS_ALIASThis domain is an alias: every path redirects to another sitelow0.7%rev 1
- Why it matters
- Every request to this hostname is redirected to a different site, so nothing measured here describes a site at this address. Scan the destination instead; the redirect itself is the only fact about this domain.
- Fix
- Nothing to fix on this domain: it forwards to another site. Scan the destination host instead.
- Share of scans
- 0.7% — 21 of 3137 counted scans
- Links
- JSON · glossary: canonical · glossary: canonical-url · glossary: alias-domain
HOMEPAGE_REDIRECTS_OFF_HOSTThe homepage sends crawlers to a different hostnamemedium2%rev 1
- Why it matters
- A request for this domain's homepage answers with a redirect to a hostname outside its own www/apex pair - a regional storefront, a checkout host, a vendor platform. The scanner does not follow it, and neither does a crawler attribute what it finds there to this domain: content, schema and entity signals on the other host belong to the other host, and this domain reads as a site with no homepage. If the redirect is deliberate, the domain still needs a page of its own that a crawler can read.
- Fix
- Point the domain at the site it is meant to serve, or make the redirect a single permanent hop to the canonical host.
effort: config - Share of scans
- 2% — 62 of 3137 counted scans
- Links
- JSON · glossary: off-host-redirect
ORIGIN_BLOCKED_BEHIND_CACHEA security layer refused this file — source unconfirmedmedium0.2%rev 1
- Why it matters
- The cached copy is fine; behind it the origin refused this request. IMPORTANT LIMIT: a scanner runs from one IP, and application firewalls block by IP reputation — so this may mean the site refuses crawlers, or it may only mean the site refuses THIS scanner. The two are indistinguishable from outside. Confirm in the server access logs, grouped by source IP, before acting on it.
- Fix
- The cached copy works but the origin refuses. Fix the origin block before the cache expires, or the next miss serves the block to everyone.
effort: config - Share of scans
- 0.2% — 7 of 3137 counted scans
- Links
- JSON · glossary: edge-cache · glossary: cache-buster
STAGING_LEAKA staging hostname is exposed in the page sourcelow0.1%rev 1
- Why it matters
- A staging hostname appears in the page source. If that hostname is reachable and indexable it competes with the production site for the same content.
- Fix
- Remove staging or development hostnames, paths and comments from the production HTML, and block the staging site from crawlers.
effort: config - Share of scans
- 0.1% — 3 of 3137 counted scans
- Links
- JSON · glossary: staging-leak
Headers
HSTS_MISSINGHTTPS is served without a Strict-Transport-Security headerinfo9.3%rev 1
- Why it matters
- Strict-Transport-Security tells a browser that has visited once to use HTTPS for every later request without trying HTTP first. Without it, a typed or linked http:// address makes one plaintext request before the redirect. It does not change what crawlers receive or how the site ranks; it is a small, reversible hardening step, and it is listed at the lowest severity for that reason.
- Fix
- Send Strict-Transport-Security: max-age=15552000; includeSubDomains on HTTPS responses (Cloudflare: SSL/TLS, Edge Certificates, HSTS). Leave preload off unless you intend a change that takes months to undo.
generate the fix
effort: config - Share of scans
- 9.3% — 292 of 3137 counted scans
- Links
- JSON · glossary: hsts
HSTS_PRELOAD_INELIGIBLEThe HSTS header asks for browser preloading but does not qualify for itlow4.1%rev 1
- Why it matters
- The header carries the `preload` token, which is a request to be compiled into the HTTPS-only list shipped inside browsers. That list only accepts a header with max-age of at least 31536000 (one year) AND includeSubDomains AND preload. As sent, the token is ignored: the site is declaring an intent it does not meet, and if it ever was on the list, this header is what gets it removed. The rest of the HSTS header still works.
- Fix
- Either raise max-age to 31536000 and keep includeSubDomains so the preload token qualifies, or remove the preload token. As sent it does nothing.
generate the fix
effort: config - Share of scans
- 4.1% — 130 of 3137 counted scans
- Links
- JSON · glossary: hsts · glossary: hsts-preload-list
HSTS_PRELOAD_REMOVEDThis domain was removed from the browser HSTS preload listlow0.4%rev 1
- Why it matters
- The browser preload list (hstspreload.org) records this domain as removed. Domains are removed when the header stops meeting the requirements, or on request. Browsers that shipped with the old list still force HTTPS on every subdomain; newer ones do not, so behaviour now differs by browser version. Either the removal was intended and the `preload` token should go, or it was not and the header needs to qualify again before resubmitting.
- Fix
- If preloading was never intended, remove the preload token. If it was, restore a qualifying header (max-age at least 31536000, includeSubDomains, preload) and resubmit at hstspreload.org.
effort: config - Share of scans
- 0.4% — 14 of 3137 counted scans
- Links
- JSON · glossary: hsts · glossary: hsts-preload-list
STALE_CACHE_SERVEDVisitors and crawlers are being served an old copy of this pagelow17.5%rev 2 · 2026-09-30
- Why it matters
- The page handed to us carried a cache age older than an hour. The detail says whether the copy still matches what the origin produces now: if it matches, nothing is broken today; if it differs, an edit is already waiting behind the cache. It matters the moment you change something: an edit, a new price, a corrected phone number stays invisible to every visitor and every crawler for about that long, and nothing in your dashboard says so.
- Fix
- Lower the cache lifetime on HTML (an hour or less), or purge the edge and origin caches after every publish so crawlers see the current copy.
generate the fix
effort: config - Share of scans
- 17.5% — 549 of 3137 counted scans
- Links
- JSON · glossary: stale-response · glossary: edge-cache · glossary: age-header · the fix end to end on a real site
Mostly code
CONTENT_NEEDS_JAVASCRIPTThe page’s content only exists after JavaScript runsmedium1.1%rev 1
- Why it matters
- The HTML the server sends is an application shell: a framework mount point and a script bundle, with almost no readable text. A browser runs the bundle and the page appears. A crawler that does not execute JavaScript — which is most AI crawlers and every training-data fetch — receives the shell and quotes nothing. This is measured on the bytes the server delivered, not on a rendered view; a framework by itself is not the defect, an empty delivery is.
- Fix
- Render the page on the server or at build time (Next.js/Nuxt/SvelteKit SSR or SSG, or a prerender step) so the delivered HTML already contains the text. Check by fetching the page with curl: the words a crawler needs must be in that response.
effort: build - Share of scans
- 1.1% — 34 of 3137 counted scans
- Links
- JSON · glossary: render-dependence · glossary: render-dependence-gap · glossary: render-gap
PAGE_IS_MOSTLY_CODEAlmost nothing a machine receives from this page is readable textmedium29.2%rev 1
- Why it matters
- A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.
- Fix
- Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.
generate the fix
effort: build - Share of scans
- 29.2% — 916 of 3137 counted scans
- Links
- JSON · glossary: payload-text-share · the fix end to end on a real site
PLACEHOLDER_CONTENTUnfinished placeholder text is live on this pagehigh0.2%rev 1
- Why it matters
- Template text that was never replaced is being served to visitors and to crawlers. An answer engine cannot tell draft copy from real copy — it reads and may quote whatever is there. Placeholder in a title or meta description is worse than in body copy, because that is the text a search result shows and an assistant repeats.
- Fix
- Replace the placeholder text (lorem ipsum, 'test content', template copy) with the real page, or unpublish the page until it is written.
effort: content - Share of scans
- 0.2% — 5 of 3137 counted scans
- Links
- JSON · glossary: placeholder-content
Machine files
SITEMAP_NO_LASTMODThe sitemap does not say when anything changedlow6.6%rev 1
- Why it matters
- The sitemap lists URLs but carries no lastmod dates, so it says what exists and not what changed. A crawler with a budget uses lastmod to decide what to re-fetch; without it, the choice is made for you — usually by re-crawling little and late. This is the change signal every crawler already understands, and unlike a push protocol it needs no key, no account and no participation from the search engine.
- Fix
- Add <lastmod> dates to sitemap entries so a crawler can prioritise pages that changed.
generate the fix
effort: content - Share of scans
- 6.6% — 206 of 3137 counted scans
- Links
- JSON · glossary: sitemap · glossary: sitemap-lastmod
MACHINE_CHAIN_HEAVYThe machine files cost more than the pageinfo17.8%rev 2 · 2026-09-30
- Why it matters
- A crawler reads robots.txt, the sitemap, llms.txt and the agent files before the first page, and pays for every 404 that answers with a full page. Keep machine files as small as their job allows and serve absent paths with a short 404 body.
- Fix
- generate the fix
- Share of scans
- 17.8% — 559 of 3137 counted scans
- Links
- JSON · glossary: machine-layer
MACHINE_FILE_CACHE_SPLITThis file answers differently depending on whether the cache is hithigh0.1%rev 1
- Fix
- Make the apex and www hosts serve the same machine files, or redirect one host to the other so there is only one copy to keep current.
effort: config - Share of scans
- 0.1% — 2 of 3137 counted scans
- Links
- JSON · glossary: edge-cache-pinning · glossary: cache-split · glossary: cache-buster
MACHINE_FILE_CACHE_STALEThe cache is still serving a file the origin no longer haslow0.6%rev 1
- Fix
- Give robots.txt, sitemap.xml and llms.txt a short cache lifetime (five minutes is plenty) so an edit reaches crawlers the same day.
effort: config - Share of scans
- 0.6% — 18 of 3137 counted scans
- Links
- JSON · glossary: cache-split · glossary: edge-cache · glossary: cache-buster
SITEMAP_IS_HTMLThe sitemap returns HTML, not XMLcritical0.7%rev 1
- Why it matters
- A crawler that cannot parse the sitemap cannot enumerate the site. It falls back to following links, so deep and newly published pages go undiscovered.
- Fix
- The sitemap URL is returning an HTML page. Serve XML at that path, or update robots.txt to point at the sitemap that actually exists.
effort: config - Share of scans
- 0.7% — 23 of 3137 counted scans
- Links
- JSON · glossary: sitemap · glossary: soft-404-machine-file
SITEMAP_BLOCKEDThe sitemap is refusedhigh2.4%rev 1
- Why it matters
- SAME LIMIT AS ABOVE: if a named application firewall appears in the response, this may be blocking the scanner's IP rather than crawlers. Confirm in server logs grouped by source IP. Otherwise: the server refuses this sitemap rather than returning it or a clean 404. Crawlers cannot enumerate the site and fall back to link-following, so deep and newly published pages go undiscovered.
- Fix
- Allow the sitemap URL through the security layer so crawlers can read it.
effort: config - Share of scans
- 2.4% — 74 of 3137 counted scans
- Links
- JSON · glossary: sitemap
SITEMAP_ORIGIN_ERRORThe origin returned a server error on the sitemap pathlow0.2%rev 1
- Fix
- The sitemap URL returned a server error. Regenerate it (most CMSs have a setting) and check the URL loads in a private window.
effort: config - Share of scans
- 0.2% — 6 of 3137 counted scans
- Links
- JSON · glossary: sitemap · glossary: sitemap-origin-error
SITEMAP_UNREADABLEWhether a sitemap exists could not be determined — every path was refusedlow0.5%rev 1
- Why it matters
- A refusal is not an absence. The paths where a sitemap would live answered with a block, so this scan cannot say whether one exists — and it does not guess. Confirm from a second network, and check server logs grouped by source IP: if the block is aimed at this scanner's address, the finding is about our access; if named crawlers see the same refusal, the sitemap is unreachable to them too.
- Fix
- Fix the sitemap XML so it parses: one <urlset> or <sitemapindex>, valid <loc> entries, no HTML wrapper.
effort: build - Share of scans
- 0.5% — 16 of 3137 counted scans
- Links
- JSON · glossary: sitemap · glossary: sitemap-index
SITEMAP_DECLARES_NOTHINGA sitemap answers 200 and lists nothinglow1.3%rev 1
- Why it matters
- The file parses as a sitemap and contains no URLs. A crawler that follows the declaration fetches it on every visit and learns nothing, and Search Console reports 0 discovered URLs for it. The usual cause is a generator with an empty source — an entity-map, image or news sitemap whose feed returned nothing — or a rewrite rule left behind. Either populate it or stop declaring it.
- Fix
- Populate the empty sitemap or remove it from robots.txt and from the sitemap index. A file that lists nothing should not be declared.
effort: build - Share of scans
- 1.3% — 40 of 3137 counted scans
- Links
- JSON · glossary: sitemap
NO_SITEMAP_FOUNDNo XML sitemap foundhigh8.7%rev 1
- Why it matters
- Without a sitemap, discovery depends entirely on internal linking. Orphaned and recently published pages may never be found.
- Fix
- Publish a sitemap at /sitemap.xml (most CMSs generate one) and name it in robots.txt.
effort: build - Share of scans
- 8.7% — 272 of 3137 counted scans
- Links
- JSON · glossary: sitemap
NO_LLMS_TXTNo llms.txtinfo30.9%rev 1
- Why it matters
- Not a defect. llms.txt is an emerging convention for telling AI systems what a site is and which pages matter.
- Fix
- Publish /llms.txt: a short Markdown file naming the business, what it does, and the pages worth reading, served as text/plain or text/markdown.
generate the fix
effort: content - Share of scans
- 30.9% — 970 of 3137 counted scans
- Links
- JSON · glossary: llms-txt · glossary: machine-layer · the fix end to end on a real site
Robots
ROBOTS_RULES_SHADOWEDrobots.txt rules do not apply to named agentsmedium13.4%rev 2 · 2026-09-30
- Why it matters
- A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.
- Fix
- Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.
generate the fix
effort: config - Share of scans
- 13.4% — 419 of 3137 counted scans
- Links
- JSON · glossary: robots-txt · glossary: robots-group-merging
INSTRUCTION_CONFLICTThe instruction files contradict each otherlow0.3%rev 1
- Why it matters
- A crawler obeys only its own User-agent group. A Content-Signal or Disallow written under * says nothing to an agent that has its own group, and a group that both invites and forbids leaves the crawler to pick. Put the signal in every group it is meant for, and let the rules under it agree.
- Fix
- the finding in your report names what to change
- Share of scans
- 0.3% — 8 of 3137 counted scans
- Links
- JSON · glossary: x-robots-tag
ROBOTS_IS_HTMLrobots.txt is HTML, not textcritical0.2%rev 1
- Why it matters
- robots.txt must be plain text. Served as HTML it cannot be parsed, so no crawl directives and no sitemap reference are read.
- Fix
- robots.txt is returning an HTML page. Serve a real text file at /robots.txt, and check that no catch-all route is answering instead.
effort: config - Share of scans
- 0.2% — 6 of 3137 counted scans
- Links
- JSON · glossary: robots-txt · glossary: robots-parse-error · glossary: soft-404-machine-file
ROBOTS_IS_CATCHALLrobots.txt answers with the catch-all page - there is no robots.txtlow0.6%rev 1
- Why it matters
- The site answers robots.txt with the same page it serves for a path that does not exist. Crawlers cannot parse HTML as robots rules, so they proceed as if no file existed: everything allowed. Publish a real text/plain robots.txt.
- Fix
- A catch-all route returns the homepage for any path, including robots.txt. Add a real robots.txt and return 404 for paths that do not exist.
effort: config - Share of scans
- 0.6% — 20 of 3137 counted scans
- Links
- JSON · glossary: robots-txt
ROBOTS_BLOCKEDrobots.txt is unreadablehigh0.5%rev 1
- Why it matters
- A 5xx on robots.txt tells a standards-following crawler to treat the whole site as disallowed (RFC 9309). A 401 or 403 is a 4xx, which the standard treats as no file and so allows everything, though some crawlers read it as a refusal. Either way the rules you wrote are not the rules a crawler read. This is the single most expensive file on the site to get wrong.
- Fix
- robots.txt could not be fetched. Allow it through the security layer; it is the one file every crawler reads first.
effort: config - Share of scans
- 0.5% — 15 of 3137 counted scans
- Links
- JSON · glossary: robots-txt · glossary: robots-unavailable
ROBOTS_DISALLOW_ALLrobots.txt blocks every crawler it does not namecritical1.5%rev 2 · 2026-09-30
- Why it matters
- robots.txt instructs every crawler to stay off the entire site.
- Fix
- robots.txt disallows everything. Remove the blanket Disallow: / rule (or restrict it to the paths you actually mean).
effort: config - Share of scans
- 1.5% — 48 of 3137 counted scans
- Links
- JSON · glossary: robots-txt
DECLARED_SITEMAP_REFUSEDThe sitemap robots.txt declares refused this clientlow0.5%rev 1
- Fix
- The sitemap named in robots.txt refuses crawlers. Allow it through the security layer.
effort: config - Share of scans
- 0.5% — 17 of 3137 counted scans
- Links
- JSON · glossary: sitemap
DECLARED_SITEMAP_BROKENrobots.txt points at a sitemap that failshigh1.3%rev 1
- Why it matters
- robots.txt advertises a sitemap URL that does not resolve. Crawlers that trust the declaration and fail get nothing.
- Fix
- robots.txt names a sitemap that does not resolve. Fix the URL in robots.txt or publish the sitemap at the path it names.
effort: build - Share of scans
- 1.3% — 42 of 3137 counted scans
- Links
- JSON · glossary: sitemap · glossary: sitemap-index
ROBOTS_NOT_200robots.txt does not return 200medium3.9%rev 1
- Why it matters
- The response code on robots.txt determines how crawlers treat the whole site. Anything other than 200 or a clean 404 is risky.
- Fix
- Serve robots.txt with status 200 as text/plain. A 404 or 410 means no file: under RFC 9309 a crawler may then fetch everything, so any rule you meant to publish is not in force. A 429 is read by Google as a server error, which pauses crawling: exempt robots.txt from rate limits.
effort: config - Share of scans
- 3.9% — 123 of 3137 counted scans
- Links
- JSON · glossary: robots-txt · glossary: robots-unavailable
ROBOTS_NO_SITEMAProbots.txt declares no sitemapmedium9%rev 1
- Why it matters
- Adding one Sitemap: line lets crawlers enumerate the site instead of guessing at it.
- Fix
- Add a Sitemap: line to robots.txt pointing at the full sitemap URL, so crawlers find it without guessing.
generate the fix
effort: config - Share of scans
- 9% — 281 of 3137 counted scans
- Links
- JSON · glossary: robots-txt
AI_OPTOUT_SETAn AI opt-out signal is setmedium5%rev 1
- Why it matters
- A Content-Signal directive is telling AI systems not to use this site. CDNs inject this by default on some plans, so it is frequently set without the owner deciding to set it.
- Fix
- This is a choice, not a defect. If you want answer engines to quote the site, remove the AI opt-out directives; if you do not, leave it.
effort: config - Share of scans
- 5% — 156 of 3137 counted scans
- Links
- JSON · glossary: ai-opt-out · glossary: machine-readable-opt-out
ROBOTS_CTYPErobots.txt has an unexpected content typelow0.1%rev 1
- Why it matters
- robots.txt should be served as text/plain. Some parsers are strict about this.
- Fix
- Serve robots.txt as text/plain. Some validators refuse any other content type.
effort: config - Share of scans
- 0.1% — 3 of 3137 counted scans
- Links
- JSON · glossary: robots-txt · glossary: robots-parse-error
RENDER_RESOURCE_BLOCKEDThe homepage depends on files this site's own robots.txt disallowsmedium2.6%rev 2 · 2026-09-30
- Why it matters
- A crawler that obeys robots.txt cannot load those files, so it renders a different page from the one visitors see.
- Fix
- Allow crawlers to fetch the CSS and JavaScript the page depends on. Remove the Disallow rules covering those paths, or serve the content without them.
effort: config - Share of scans
- 2.6% — 83 of 3137 counted scans
- Links
- JSON · glossary: blocked-render-resource
Delivery parity
CHALLENGE_SERVED_200A verification page is being served instead of the filecritical0.6%rev 1
- Why it matters
- The server answered with HTTP 200 — so uptime monitors, Search Console and your host all report this as healthy — but the body is a JavaScript verification page, not the file. Crawlers do not run JavaScript on these requests. They receive the interstitial and move on.
- Fix
- A verification page must not answer with status 200. Return 403 or 503 for challenges so crawlers do not index the challenge as the page.
effort: config - Share of scans
- 0.6% — 20 of 3137 counted scans
- Links
- JSON · glossary: challenge-page-at-200
ANSWER_ENGINE_REFUSEDAn answer engine's crawler was refused while a browser was servedhigh5.9%rev 1
- Why it matters
- The same URL, the same minute: an ordinary client was served the page and a named answer-engine crawler was not. That is the edge rather than a robots.txt decision. One caveat this site's own guides state and this finding must too: the scan sent the crawler's name from an address that operator does not publish, and an edge that verifies identity by IP is right to refuse it. So this proves the edge treats the identity differently - not that it refuses the real crawler. Your own logs, checked against the operator's published ranges, settle which. If the edge is refusing the real crawler, nothing else in this report reaches that engine, however good it is.
- Fix
- Allow the named answer-engine crawlers through the security layer (Cloudflare bot rules, WAF, rate limits), or verify their published IP ranges instead of blocking by name.
effort: config - Share of scans
- 5.9% — 184 of 3137 counted scans
- Links
- JSON · glossary: refusal · glossary: delivery-comparison
UNIFORM_REFUSALThis origin refused the scanner, so the site was never measuredcritical3.6%rev 1
- Fix
- The site refused every request from this scanner. Allow this scanner's address, or run the browser extension from your own network, then re-scan.
effort: config - Share of scans
- 3.6% — 112 of 3137 counted scans
- Links
- JSON · glossary: uniform-refusal · glossary: datacenter-ip
CRAWLER_SERVED_LESSA named crawler was served materially less text than an ordinary clienthigh0.4%rev 2 · 2026-09-30
- Why it matters
- Same URL, same minute, both answered 200 — but the named crawler received materially fewer words than an unnamed client did. A refusal is loud and easy to find; being served a thinner page is silent, and it is the version an answer engine quotes from.
- Fix
- Serve every crawler the same HTML a browser gets. Check bot-management rules, 'lite' templates and user-agent conditions in the CMS or CDN.
effort: config - Share of scans
- 0.4% — 14 of 3137 counted scans
- Links
- JSON · glossary: delivery-comparison
CRAWLER_REDIRECTED_AWAYA named crawler was sent somewhere the ordinary client was nothigh0.4%rev 2 · 2026-09-30
- Why it matters
- The ordinary client stayed on the URL and the named crawler was sent elsewhere. Whatever is at the destination is what that engine will read and quote instead of this page.
- Fix
- Remove the user-agent based redirect so named crawlers land on the page a visitor lands on.
effort: config - Share of scans
- 0.4% — 11 of 3137 counted scans
- Links
- JSON · glossary: delivery-comparison
CHALLENGE_PINNED_AT_EDGEThe verification page is cached at the CDN edgecritical0.5%rev 1
- Why it matters
- The origin marked this response private and uncacheable. A cache rule with a TTL override stored it anyway, so the verification page is now being handed to every visitor and every crawler from the edge until it expires or is purged.
- Fix
- Purge the edge cache: a challenge page was cached in place of the real file. Then exclude machine files from HTML caching rules.
effort: config - Share of scans
- 0.5% — 15 of 3137 counted scans
- Links
- JSON · glossary: edge-cache-pinning · glossary: challenge-page-at-200
HOMEPAGE_REFUSEDThe homepage refused this client, so every homepage section is unmeasuredlow2%rev 1
- Fix
- The homepage was refused. Check bot-protection and WAF rules; a challenge page served with a 200 status reads as the site to a crawler.
effort: config - Share of scans
- 2% — 64 of 3137 counted scans
- Links
- JSON · glossary: homepage-refused · glossary: refusal
HOMEPAGE_IS_INTERSTITIALThe homepage answered with an interstitial, not a pagecritical0.2%rev 1
- Fix
- Remove or lower the interstitial (age gate, consent wall, JS challenge) for crawlers, or make sure the real content is in the first response.
effort: config - Share of scans
- 0.2% — 5 of 3137 counted scans
- Links
- JSON · glossary: 202-interstitial
JSON-LD and entities
ENTITY_COLLISIONTwo different entities share the same identifying datamedium3.5%rev 1
- Why it matters
- Two entities on the page resolve to the same identifying data (e.g. name, phone, or place ID), so they cannot both be correct at once.
- Fix
- Merge the duplicate business nodes into one, keep a single @id, and remove any plugin that emits a second empty LocalBusiness record.
generate the fix
effort: build - Share of scans
- 3.5% — 110 of 3137 counted scans
- Links
- JSON · glossary: entity-collision
ENTITY_NO_COORDINATESAn entity is missing geographic coordinatesinfo5.7%rev 1
- Why it matters
- An entity is missing latitude/longitude coordinates, so it cannot be verified against a real-world location.
- Fix
- Add latitude and longitude to the LocalBusiness node's geo property so an engine can place the business on a map without guessing.
generate the fix
effort: content - Share of scans
- 5.7% — 178 of 3137 counted scans
- Links
- JSON · glossary: geocoordinates · glossary: geocoding · glossary: postaladdress
GEO_COUNTRY_MISMATCHPublished coordinates are not in the country the address nameslow0.2%rev 1
- Why it matters
- The latitude/longitude on the page falls outside the country its own address names. A range check cannot catch this: a missing minus sign is still a legal coordinate. Coordinates are read only by machines, so the error produces no visible symptom.
- Fix
- the finding in your report names what to change
- Share of scans
- 0.2% — 6 of 3137 counted scans
- Links
- JSON · glossary: geocoordinates · glossary: geocoding · glossary: postaladdress
NAP and claims
NAP_NAME_DRIFTA directory prints a different business name than the site declareslow0.7%rev 1
- Why it matters
- Two names on one phone number read as two entities, or as one entity with an unsettled name; an engine reconciling them may print either.
- Fix
- the finding in your report names what to change
- Share of scans
- 0.7% — 21 of 3137 counted scans
- Links
- JSON · glossary: nap · glossary: nap-consistency · glossary: name-drift
NAP_UNDECLARED_LISTINGOff-site listings carry this phone number and the site does not declare theminfo4.9%rev 1
- Why it matters
- A listing the site does not declare is one an engine corroborates without you: whatever it prints becomes your name, hours and category in the answer.
- Fix
- the finding in your report names what to change
- Share of scans
- 4.9% — 153 of 3137 counted scans
- Links
- JSON · glossary: nap · glossary: nap-consistency
Answer-engine citations
CITED_WITH_WRONG_FACTSAn answer engine cites the site with the wrong phone or addressmedium—rev 1
- Why it matters
- The engine read the business facts from somewhere other than the site, usually a directory, and is repeating them.
- Fix
- Fix the source the engine read (usually a directory listing), then make the site's schema and sameAs agree with it.
- Share of scans
- measured in answer runs, not counted per scan
- Links
- JSON
REACHABLE_NOT_CITEDAn answer engine can reach the site and does not cite itinfo—rev 1
- Why it matters
- Access is not the problem: the content and authority are.
- Fix
- See the Quote pillar rows on the machine-layer report: defining lead, quotable paragraphs, FAQ parity.
- Share of scans
- measured in answer runs, not counted per scan
- Links
- JSON
BOT_BLOCKED_NOT_CITEDAn answer engine does not cite the site and robots.txt blocks its crawlermedium—rev 1
- Why it matters
- An engine cannot quote a page its crawler is told not to read.
- Fix
- Allow the engine's crawler in robots.txt (User-agent: <crawler> / Allow: /).
- Share of scans
- measured in answer runs, not counted per scan
- Links
- JSON
BOT_RULES_SHADOWED_NOT_CITEDAn answer engine does not cite the site and its crawler group is shadowedlow—rev 1
- Why it matters
- The crawler's own group in robots.txt does not carry the rules the site meant it to follow.
- Fix
- Move the named crawler group above the wildcard group, or repeat its Allow lines after it.
- Share of scans
- measured in answer runs, not counted per scan
- Links
- JSON
NO_LLMS_TXT_NOT_CITEDAn answer engine does not cite the site and there is no llms.txtinfo—rev 1
- Why it matters
- The engine can reach the site; the site gives it no summary to start from.
- Fix
- Publish /llms.txt (the scanner's /api/tool/llms drafts one from the sitemap).
- Share of scans
- measured in answer runs, not counted per scan
- Links
- JSON
Capabilities
CAPABILITY_DECLARED_PUBLIC_OBSERVED_AUTHDeclared capability declared public observed authnot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
CAPABILITY_DECLARED_ENDPOINT_MISSINGDeclared capability declared endpoint missingnot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
CAPABILITY_DECLARED_NO_REQUIRED_INPUT_OBSERVED_REJECTEDDeclared capability declared no required input observed rejectednot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
CAPABILITY_DECLARED_MEDIA_MISMATCHDeclared capability declared media mismatchnot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
CAPABILITY_DECLARED_TOOL_NOT_LISTEDDeclared capability declared tool not listednot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
CAPABILITY_LISTED_TOOL_NOT_DECLAREDDeclared capability listed tool not declarednot scored—rev 1
- Fix
- the finding in your report names what to change
- Share of scans
- listed in the record's capabilities section, not counted per scan
- Links
- JSON
Self-audit: defects in our own record
FETCHER_DISAGREEMENTTwo of our fetch layers disagree about a machine filemedium—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
BENCH_POOL_MISMATCHThe benchmark's pool size contradicts its scopemedium—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
COHORT_BELOW_THRESHOLDRanked against a cohort under the minimumlow—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
SECTION_SCORE_OUT_OF_RANGEA section score outside 0 to 100high—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
RANKED_VALUE_NOT_REPRODUCIBLEThe ranked number does not follow from the printed tablehigh—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
SECTION_WEIGHT_UNDECLAREDA section scored without a declared weightmedium—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
PATH_NOT_MEASUREDA deep path was asked for, the home page was measuredlow—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
PLACEHOLDER_IN_RECORDOne of our placeholders reached the recordhigh—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
GRADE_OFF_ITS_OWN_SCALEThe letter does not follow the record's own thresholdshigh—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
GRADED_THROUGH_A_WALLA grade on a site that refused the scannercritical—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- The reading was taken through a block page. Allow the scanner and re-scan for a real grade.
effort: config - Share of scans
- about our own record, not counted per scan
- Links
- JSON
EXTRACTORS_READ_NOTHINGOur extractors read nothing from a head that has tagsmedium—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
GRADE_MISSINGA graded record with no grademedium—rev 1
- Why it matters
- A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
- Fix
- the finding in your report names what to change
- Share of scans
- about our own record, not counted per scan
- Links
- JSON
MCP servers
MCP_TOOL_POISONINGAn MCP server's tool descriptions carry text aimed at the modelhigh—rev 1
- Why it matters
- Text a client hands the model from tools/list hides characters a person cannot see, tells the model to ignore its instructions or keep something from the user, asks for secrets or the conversation, or tells the model how to use other tools. A model reads it as an instruction; the person approving the server usually never sees it.
- Fix
- Remove the text from the tool name, description and schema; descriptions should say what the tool does and nothing to the model about other tools, secrets or the user. Re-check with /api/v1/mcp/scan.
- Share of scans
- measured on remote MCP servers in the official MCP Registry, not per site scan: see https://crawlcheck.io/api/mcp/index (tool_poisoning)
- Links
- JSON
High-risk field watch
HIGH_RISK_FIELD_CHANGEDA high-risk field changed and the owner has not confirmed ithigh—rev 1
- Why it matters
- One of the fields an attacker changes first moved since CrawlCheck last read it: the payment processors loaded on the homepage or checkout page, the host checkout links go to, a cryptocurrency wallet address printed on the page, the OAuth authorization server the site's MCP endpoint names, the domain's MX hosts, or its registry-listed MCP endpoints. Watched fields: payment_processors, checkout_host, wallet_addresses, oauth_authorization_server, mx_hosts, mcp_endpoints. The first reading is the baseline and never alerts.
- Fix
- If the change is yours, confirm it: POST /api/v1/high-risk/confirm {domain} with the key that holds the domain claim. If it is not yours, treat the site as compromised: flip the incident switch (POST /api/v1/incident) and restore the field.
- Share of scans
- measured by the daily high-risk field watch, not per site scan: see https://crawlcheck.io/api/v1/high-risk
- Links
- JSON
Generated 2026-10-07 16:48 UTC. A rule you think is wrong? Write to hello@crawlcheck.io with the report id; corrections are published as a new revision.