{"ok":true,"kind":"crawlcheck-rulebook","v":1,"generated_at":"2026-10-07T17:43:46.602Z","score_version":28,"schema":"https://crawlcheck.io/schemas/rulebook.json","counts":{"rules":77,"by_kind":{"scan":52,"answer_correlation":5,"capability_mismatch":6,"self_audit":12,"mcp_server":2},"by_family":{"site":8,"headers":4,"mostly_code":3,"machine_files":11,"robots":13,"delivery":8,"jsonld":3,"nap":2,"answers":5,"capabilities":6,"self_audit":12,"mcp":1,"watch":1},"revised":7,"scored":48},"scans_counted":3137,"severity_scale":[{"level":"info"},{"level":"low"},{"level":"medium"},{"level":"high"},{"level":"critical"}],"method":{"revisions":"A rule that changes what it reports gets a new revision; a record keeps the revision that decided it, and a replay uses that revision, never today's.","share":"Share of counted scans that carried the code at least once, the same numbers /data publishes. Self-scans and opted-out sites are never counted.","grade":"Each finding has a level from info to critical. Open findings at medium and above lower the letter grade; fix them first."},"rules":[{"code":"CSP_REPORTING_BROKEN","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"CSP violation reporting is misconfigured","meaning":"The Content-Security-Policy header declares where to send violation reports, and the declaration is wrong. `report-to` takes the NAME of a group defined in a separate header — given a URL it is silently ignored, so nothing is reported at all. And `report-uri` aimed at your own origin turns every violation into another request to the page you are already trying to keep fast. Neither breaks the page; both mean the reporting you think you have is either absent or self-inflicted.","severity":{"level":"info"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The Content-Security-Policy report endpoint does not accept reports. Fix the report-uri / report-to target or remove the directive.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.4,"scans":11,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CSP_REPORTING_BROKEN","api":"https://crawlcheck.io/api/rules?code=CSP_REPORTING_BROKEN","glossary":["https://crawlcheck.io/glossary/content-security-policy"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CSP_REPORTING_BROKEN"}},{"code":"HSTS_MISSING","kind":"scan","family":{"id":"headers","label":"Headers"},"title":"HTTPS is served without a Strict-Transport-Security header","meaning":"Strict-Transport-Security tells a browser that has visited once to use HTTPS for every later request without trying HTTP first. Without it, a typed or linked http:// address makes one plaintext request before the redirect. It does not change what crawlers receive or how the site ranks; it is a small, reversible hardening step, and it is listed at the lowest severity for that reason.","severity":{"level":"info"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Send Strict-Transport-Security: max-age=15552000; includeSubDomains on HTTPS responses (Cloudflare: SSL/TLS, Edge Certificates, HSTS). Leave preload off unless you intend a change that takes months to undo.","effort":"config","endpoint":"https://crawlcheck.io/api/fix/hsts?domain={domain}"},"share_of_scans":{"pct":9.3,"scans":292,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HSTS_MISSING","api":"https://crawlcheck.io/api/rules?code=HSTS_MISSING","glossary":["https://crawlcheck.io/glossary/hsts"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HSTS_MISSING"}},{"code":"HSTS_PRELOAD_INELIGIBLE","kind":"scan","family":{"id":"headers","label":"Headers"},"title":"The HSTS header asks for browser preloading but does not qualify for it","meaning":"The header carries the `preload` token, which is a request to be compiled into the HTTPS-only list shipped inside browsers. That list only accepts a header with max-age of at least 31536000 (one year) AND includeSubDomains AND preload. As sent, the token is ignored: the site is declaring an intent it does not meet, and if it ever was on the list, this header is what gets it removed. The rest of the HSTS header still works.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Either raise max-age to 31536000 and keep includeSubDomains so the preload token qualifies, or remove the preload token. As sent it does nothing.","effort":"config","endpoint":"https://crawlcheck.io/api/fix/hsts?domain={domain}"},"share_of_scans":{"pct":4.1,"scans":130,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HSTS_PRELOAD_INELIGIBLE","api":"https://crawlcheck.io/api/rules?code=HSTS_PRELOAD_INELIGIBLE","glossary":["https://crawlcheck.io/glossary/hsts","https://crawlcheck.io/glossary/hsts-preload-list"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HSTS_PRELOAD_INELIGIBLE"}},{"code":"HSTS_PRELOAD_REMOVED","kind":"scan","family":{"id":"headers","label":"Headers"},"title":"This domain was removed from the browser HSTS preload list","meaning":"The browser preload list (hstspreload.org) records this domain as removed. Domains are removed when the header stops meeting the requirements, or on request. Browsers that shipped with the old list still force HTTPS on every subdomain; newer ones do not, so behaviour now differs by browser version. Either the removal was intended and the `preload` token should go, or it was not and the header needs to qualify again before resubmitting.","severity":{"level":"low"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"If preloading was never intended, remove the preload token. If it was, restore a qualifying header (max-age at least 31536000, includeSubDomains, preload) and resubmit at hstspreload.org.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.4,"scans":14,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HSTS_PRELOAD_REMOVED","api":"https://crawlcheck.io/api/rules?code=HSTS_PRELOAD_REMOVED","glossary":["https://crawlcheck.io/glossary/hsts","https://crawlcheck.io/glossary/hsts-preload-list"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HSTS_PRELOAD_REMOVED"}},{"code":"CONTENT_NEEDS_JAVASCRIPT","kind":"scan","family":{"id":"mostly_code","label":"Mostly code"},"title":"The page’s content only exists after JavaScript runs","meaning":"The HTML the server sends is an application shell: a framework mount point and a script bundle, with almost no readable text. A browser runs the bundle and the page appears. A crawler that does not execute JavaScript — which is most AI crawlers and every training-data fetch — receives the shell and quotes nothing. This is measured on the bytes the server delivered, not on a rendered view; a framework by itself is not the defect, an empty delivery is.","severity":{"level":"medium"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Render the page on the server or at build time (Next.js/Nuxt/SvelteKit SSR or SSG, or a prerender step) so the delivered HTML already contains the text. Check by fetching the page with curl: the words a crawler needs must be in that response.","effort":"build","endpoint":null},"share_of_scans":{"pct":1.1,"scans":34,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CONTENT_NEEDS_JAVASCRIPT","api":"https://crawlcheck.io/api/rules?code=CONTENT_NEEDS_JAVASCRIPT","glossary":["https://crawlcheck.io/glossary/render-dependence","https://crawlcheck.io/glossary/render-dependence-gap","https://crawlcheck.io/glossary/render-gap"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CONTENT_NEEDS_JAVASCRIPT"}},{"code":"SITEMAP_NO_LASTMOD","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The sitemap does not say when anything changed","meaning":"The sitemap lists URLs but carries no lastmod dates, so it says what exists and not what changed. A crawler with a budget uses lastmod to decide what to re-fetch; without it, the choice is made for you — usually by re-crawling little and late. This is the change signal every crawler already understands, and unlike a push protocol it needs no key, no account and no participation from the search engine.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Add <lastmod> dates to sitemap entries so a crawler can prioritise pages that changed.","effort":"content","endpoint":"https://crawlcheck.io/api/fix/sitemap?domain={domain}"},"share_of_scans":{"pct":6.6,"scans":206,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_NO_LASTMOD","api":"https://crawlcheck.io/api/rules?code=SITEMAP_NO_LASTMOD","glossary":["https://crawlcheck.io/glossary/sitemap","https://crawlcheck.io/glossary/sitemap-lastmod"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_NO_LASTMOD"}},{"code":"BROKEN_INTERNAL_LINK","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"Links on this site point at pages that do not answer","meaning":"A link on this site was followed and the target answered 4xx or 5xx. A crawler spends budget on every one of these and learns nothing; a reader following one gets an error page. The usual causes are a route renamed without updating the links to it, and a nav or footer that renders links to authenticated-only routes for everyone — those answer 404 when signed out, which is what a crawler always is. Only 404 and 410 count here. A link answering 401, 403 or 429 means something declined to answer this scanner — bot protection, a login wall, a rate limit — which is a fact about the request, not about the link, so those are counted and never reported as broken. Links are sampled rather than exhaustively crawled, so this is a floor and not a full link audit.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Fix or remove the internal links that return 404 so crawlers and visitors are not sent to dead pages.","effort":"content","endpoint":null},"share_of_scans":{"pct":3,"scans":94,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#BROKEN_INTERNAL_LINK","api":"https://crawlcheck.io/api/rules?code=BROKEN_INTERNAL_LINK","glossary":["https://crawlcheck.io/glossary/broken-link","https://crawlcheck.io/glossary/internal-link","https://crawlcheck.io/glossary/dead-anchor"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=BROKEN_INTERNAL_LINK"}},{"code":"HOST_DUPLICATE_200","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"www and the apex both serve the site instead of one redirecting","meaning":"Both www and the bare domain answer 200 with the same pages, so every URL on this site exists at two addresses and a crawler fetches the whole site twice. A canonical tag does not prevent this: it is read after the response arrives, so the second fetch has already been spent. Nothing looks broken from a browser, which is why this survives for years. The fix is a 301 from one host to the other so there is one address per page.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Pick one host (apex or www) and redirect the other to it with a 301 so there is one copy of every page.","effort":"config","endpoint":null},"share_of_scans":{"pct":3.5,"scans":111,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HOST_DUPLICATE_200","api":"https://crawlcheck.io/api/rules?code=HOST_DUPLICATE_200","glossary":["https://crawlcheck.io/glossary/canonical","https://crawlcheck.io/glossary/canonical-url","https://crawlcheck.io/glossary/host-duplicate","https://crawlcheck.io/glossary/apex-domain"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HOST_DUPLICATE_200"}},{"code":"ROBOTS_RULES_SHADOWED","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt rules do not apply to named agents","meaning":"A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.","effort":"config","endpoint":"https://crawlcheck.io/api/fix/robots?domain={domain}"},"share_of_scans":{"pct":13.4,"scans":419,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_RULES_SHADOWED","api":"https://crawlcheck.io/api/rules?code=ROBOTS_RULES_SHADOWED","glossary":["https://crawlcheck.io/glossary/robots-txt","https://crawlcheck.io/glossary/robots-group-merging"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_RULES_SHADOWED"}},{"code":"INSTRUCTION_CONFLICT","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"The instruction files contradict each other","meaning":"A crawler obeys only its own User-agent group. A Content-Signal or Disallow written under * says nothing to an agent that has its own group, and a group that both invites and forbids leaves the crawler to pick. Put the signal in every group it is meant for, and let the rules under it agree.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":0.3,"scans":8,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#INSTRUCTION_CONFLICT","api":"https://crawlcheck.io/api/rules?code=INSTRUCTION_CONFLICT","glossary":["https://crawlcheck.io/glossary/x-robots-tag"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=INSTRUCTION_CONFLICT"}},{"code":"MACHINE_CHAIN_HEAVY","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The machine files cost more than the page","meaning":"A crawler reads robots.txt, the sitemap, llms.txt and the agent files before the first page, and pays for every 404 that answers with a full page. Keep machine files as small as their job allows and serve absent paths with a short 404 body.","severity":{"level":"info"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":null,"effort":null,"endpoint":"https://crawlcheck.io/api/fix/chain?id={report_id}"},"share_of_scans":{"pct":17.8,"scans":559,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#MACHINE_CHAIN_HEAVY","api":"https://crawlcheck.io/api/rules?code=MACHINE_CHAIN_HEAVY","glossary":["https://crawlcheck.io/glossary/machine-layer"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=MACHINE_CHAIN_HEAVY"}},{"code":"CHALLENGE_SERVED_200","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"A verification page is being served instead of the file","meaning":"The server answered with HTTP 200 — so uptime monitors, Search Console and your host all report this as healthy — but the body is a JavaScript verification page, not the file. Crawlers do not run JavaScript on these requests. They receive the interstitial and move on.","severity":{"level":"critical"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"A verification page must not answer with status 200. Return 403 or 503 for challenges so crawlers do not index the challenge as the page.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.6,"scans":20,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CHALLENGE_SERVED_200","api":"https://crawlcheck.io/api/rules?code=CHALLENGE_SERVED_200","glossary":["https://crawlcheck.io/glossary/challenge-page-at-200"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CHALLENGE_SERVED_200"}},{"code":"ANSWER_ENGINE_REFUSED","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"An answer engine's crawler was refused while a browser was served","meaning":"The same URL, the same minute: an ordinary client was served the page and a named answer-engine crawler was not. That is the edge rather than a robots.txt decision. One caveat this site's own guides state and this finding must too: the scan sent the crawler's name from an address that operator does not publish, and an edge that verifies identity by IP is right to refuse it. So this proves the edge treats the identity differently - not that it refuses the real crawler. Your own logs, checked against the operator's published ranges, settle which. If the edge is refusing the real crawler, nothing else in this report reaches that engine, however good it is.","severity":{"level":"high"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Allow the named answer-engine crawlers through the security layer (Cloudflare bot rules, WAF, rate limits), or verify their published IP ranges instead of blocking by name.","effort":"config","endpoint":null},"share_of_scans":{"pct":5.9,"scans":184,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ANSWER_ENGINE_REFUSED","api":"https://crawlcheck.io/api/rules?code=ANSWER_ENGINE_REFUSED","glossary":["https://crawlcheck.io/glossary/refusal","https://crawlcheck.io/glossary/delivery-comparison"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ANSWER_ENGINE_REFUSED"}},{"code":"PAGE_IS_MOSTLY_CODE","kind":"scan","family":{"id":"mostly_code","label":"Mostly code"},"title":"Almost nothing a machine receives from this page is readable text","meaning":"A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.","severity":{"level":"medium"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.","effort":"build","endpoint":"https://crawlcheck.io/api/fix/code?id={report_id}"},"share_of_scans":{"pct":29.2,"scans":916,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#PAGE_IS_MOSTLY_CODE","api":"https://crawlcheck.io/api/rules?code=PAGE_IS_MOSTLY_CODE","glossary":["https://crawlcheck.io/glossary/payload-text-share"],"workflow":"https://crawlcheck.io/workflow/PAGE_IS_MOSTLY_CODE","lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=PAGE_IS_MOSTLY_CODE"}},{"code":"STALE_CACHE_SERVED","kind":"scan","family":{"id":"headers","label":"Headers"},"title":"Visitors and crawlers are being served an old copy of this page","meaning":"The page handed to us carried a cache age older than an hour. The detail says whether the copy still matches what the origin produces now: if it matches, nothing is broken today; if it differs, an edit is already waiting behind the cache. It matters the moment you change something: an edit, a new price, a corrected phone number stays invisible to every visitor and every crawler for about that long, and nothing in your dashboard says so.","severity":{"level":"low"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"Lower the cache lifetime on HTML (an hour or less), or purge the edge and origin caches after every publish so crawlers see the current copy.","effort":"config","endpoint":"https://crawlcheck.io/api/fix/cache?domain={domain}"},"share_of_scans":{"pct":17.5,"scans":549,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#STALE_CACHE_SERVED","api":"https://crawlcheck.io/api/rules?code=STALE_CACHE_SERVED","glossary":["https://crawlcheck.io/glossary/stale-response","https://crawlcheck.io/glossary/edge-cache","https://crawlcheck.io/glossary/age-header"],"workflow":"https://crawlcheck.io/workflow/STALE_CACHE_SERVED","lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=STALE_CACHE_SERVED"}},{"code":"UNIFORM_REFUSAL","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"This origin refused the scanner, so the site was never measured","meaning":null,"severity":{"level":"critical"},"scored":false,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The site refused every request from this scanner. Allow this scanner's address, or run the browser extension from your own network, then re-scan.","effort":"config","endpoint":null},"share_of_scans":{"pct":3.6,"scans":112,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#UNIFORM_REFUSAL","api":"https://crawlcheck.io/api/rules?code=UNIFORM_REFUSAL","glossary":["https://crawlcheck.io/glossary/uniform-refusal","https://crawlcheck.io/glossary/datacenter-ip"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=UNIFORM_REFUSAL"}},{"code":"ORIGIN_REFUSED_SCANNER","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"The homepage and robots.txt both refused this client, so the site was never measured","meaning":null,"severity":{"level":"critical"},"scored":false,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Allow this scanner's address through the security layer, or scan from the browser extension, which uses your own address.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.3,"scans":8,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ORIGIN_REFUSED_SCANNER","api":"https://crawlcheck.io/api/rules?code=ORIGIN_REFUSED_SCANNER","glossary":["https://crawlcheck.io/glossary/refusal","https://crawlcheck.io/glossary/datacenter-ip","https://crawlcheck.io/glossary/vantage-point"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ORIGIN_REFUSED_SCANNER"}},{"code":"MACHINE_FILE_CACHE_SPLIT","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"This file answers differently depending on whether the cache is hit","meaning":null,"severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Make the apex and www hosts serve the same machine files, or redirect one host to the other so there is only one copy to keep current.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.1,"scans":2,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#MACHINE_FILE_CACHE_SPLIT","api":"https://crawlcheck.io/api/rules?code=MACHINE_FILE_CACHE_SPLIT","glossary":["https://crawlcheck.io/glossary/edge-cache-pinning","https://crawlcheck.io/glossary/cache-split","https://crawlcheck.io/glossary/cache-buster"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=MACHINE_FILE_CACHE_SPLIT"}},{"code":"MACHINE_FILE_CACHE_STALE","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The cache is still serving a file the origin no longer has","meaning":null,"severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Give robots.txt, sitemap.xml and llms.txt a short cache lifetime (five minutes is plenty) so an edit reaches crawlers the same day.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.6,"scans":18,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#MACHINE_FILE_CACHE_STALE","api":"https://crawlcheck.io/api/rules?code=MACHINE_FILE_CACHE_STALE","glossary":["https://crawlcheck.io/glossary/cache-split","https://crawlcheck.io/glossary/edge-cache","https://crawlcheck.io/glossary/cache-buster"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=MACHINE_FILE_CACHE_STALE"}},{"code":"CRAWLER_SERVED_LESS","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"A named crawler was served materially less text than an ordinary client","meaning":"Same URL, same minute, both answered 200 — but the named crawler received materially fewer words than an unnamed client did. A refusal is loud and easy to find; being served a thinner page is silent, and it is the version an answer engine quotes from.","severity":{"level":"high"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"Serve every crawler the same HTML a browser gets. Check bot-management rules, 'lite' templates and user-agent conditions in the CMS or CDN.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.4,"scans":14,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CRAWLER_SERVED_LESS","api":"https://crawlcheck.io/api/rules?code=CRAWLER_SERVED_LESS","glossary":["https://crawlcheck.io/glossary/delivery-comparison"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CRAWLER_SERVED_LESS"}},{"code":"CRAWLER_REDIRECTED_AWAY","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"A named crawler was sent somewhere the ordinary client was not","meaning":"The ordinary client stayed on the URL and the named crawler was sent elsewhere. Whatever is at the destination is what that engine will read and quote instead of this page.","severity":{"level":"high"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"Remove the user-agent based redirect so named crawlers land on the page a visitor lands on.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.4,"scans":11,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CRAWLER_REDIRECTED_AWAY","api":"https://crawlcheck.io/api/rules?code=CRAWLER_REDIRECTED_AWAY","glossary":["https://crawlcheck.io/glossary/delivery-comparison"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CRAWLER_REDIRECTED_AWAY"}},{"code":"CHALLENGE_PINNED_AT_EDGE","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"The verification page is cached at the CDN edge","meaning":"The origin marked this response private and uncacheable. A cache rule with a TTL override stored it anyway, so the verification page is now being handed to every visitor and every crawler from the edge until it expires or is purged.","severity":{"level":"critical"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Purge the edge cache: a challenge page was cached in place of the real file. Then exclude machine files from HTML caching rules.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.5,"scans":15,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#CHALLENGE_PINNED_AT_EDGE","api":"https://crawlcheck.io/api/rules?code=CHALLENGE_PINNED_AT_EDGE","glossary":["https://crawlcheck.io/glossary/edge-cache-pinning","https://crawlcheck.io/glossary/challenge-page-at-200"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=CHALLENGE_PINNED_AT_EDGE"}},{"code":"ROBOTS_IS_HTML","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt is HTML, not text","meaning":"robots.txt must be plain text. Served as HTML it cannot be parsed, so no crawl directives and no sitemap reference are read.","severity":{"level":"critical"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"robots.txt is returning an HTML page. Serve a real text file at /robots.txt, and check that no catch-all route is answering instead.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.2,"scans":6,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_IS_HTML","api":"https://crawlcheck.io/api/rules?code=ROBOTS_IS_HTML","glossary":["https://crawlcheck.io/glossary/robots-txt","https://crawlcheck.io/glossary/robots-parse-error","https://crawlcheck.io/glossary/soft-404-machine-file"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_IS_HTML"}},{"code":"ROBOTS_IS_CATCHALL","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt answers with the catch-all page - there is no robots.txt","meaning":"The site answers robots.txt with the same page it serves for a path that does not exist. Crawlers cannot parse HTML as robots rules, so they proceed as if no file existed: everything allowed. Publish a real text/plain robots.txt.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"A catch-all route returns the homepage for any path, including robots.txt. Add a real robots.txt and return 404 for paths that do not exist.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.6,"scans":20,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_IS_CATCHALL","api":"https://crawlcheck.io/api/rules?code=ROBOTS_IS_CATCHALL","glossary":["https://crawlcheck.io/glossary/robots-txt"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_IS_CATCHALL"}},{"code":"DOMAIN_IS_ALIAS","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"This domain is an alias: every path redirects to another site","meaning":"Every request to this hostname is redirected to a different site, so nothing measured here describes a site at this address. Scan the destination instead; the redirect itself is the only fact about this domain.","severity":{"level":"low"},"scored":false,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Nothing to fix on this domain: it forwards to another site. Scan the destination host instead.","effort":null,"endpoint":null},"share_of_scans":{"pct":0.7,"scans":21,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#DOMAIN_IS_ALIAS","api":"https://crawlcheck.io/api/rules?code=DOMAIN_IS_ALIAS","glossary":["https://crawlcheck.io/glossary/canonical","https://crawlcheck.io/glossary/canonical-url","https://crawlcheck.io/glossary/alias-domain"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=DOMAIN_IS_ALIAS"}},{"code":"ROBOTS_BLOCKED","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt is unreadable","meaning":"A 5xx on robots.txt tells a standards-following crawler to treat the whole site as disallowed (RFC 9309). A 401 or 403 is a 4xx, which the standard treats as no file and so allows everything, though some crawlers read it as a refusal. Either way the rules you wrote are not the rules a crawler read. This is the single most expensive file on the site to get wrong.","severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"robots.txt could not be fetched. Allow it through the security layer; it is the one file every crawler reads first.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.5,"scans":15,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_BLOCKED","api":"https://crawlcheck.io/api/rules?code=ROBOTS_BLOCKED","glossary":["https://crawlcheck.io/glossary/robots-txt","https://crawlcheck.io/glossary/robots-unavailable"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_BLOCKED"}},{"code":"ROBOTS_DISALLOW_ALL","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt blocks every crawler it does not name","meaning":"robots.txt instructs every crawler to stay off the entire site.","severity":{"level":"critical"},"scored":true,"measured_on":"every scan","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"robots.txt disallows everything. Remove the blanket Disallow: / rule (or restrict it to the paths you actually mean).","effort":"config","endpoint":null},"share_of_scans":{"pct":1.5,"scans":48,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_DISALLOW_ALL","api":"https://crawlcheck.io/api/rules?code=ROBOTS_DISALLOW_ALL","glossary":["https://crawlcheck.io/glossary/robots-txt"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_DISALLOW_ALL"}},{"code":"SITEMAP_IS_HTML","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The sitemap returns HTML, not XML","meaning":"A crawler that cannot parse the sitemap cannot enumerate the site. It falls back to following links, so deep and newly published pages go undiscovered.","severity":{"level":"critical"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The sitemap URL is returning an HTML page. Serve XML at that path, or update robots.txt to point at the sitemap that actually exists.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.7,"scans":23,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_IS_HTML","api":"https://crawlcheck.io/api/rules?code=SITEMAP_IS_HTML","glossary":["https://crawlcheck.io/glossary/sitemap","https://crawlcheck.io/glossary/soft-404-machine-file"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_IS_HTML"}},{"code":"SITEMAP_BLOCKED","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The sitemap is refused","meaning":"SAME LIMIT AS ABOVE: if a named application firewall appears in the response, this may be blocking the scanner's IP rather than crawlers. Confirm in server logs grouped by source IP. Otherwise: the server refuses this sitemap rather than returning it or a clean 404. Crawlers cannot enumerate the site and fall back to link-following, so deep and newly published pages go undiscovered.","severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Allow the sitemap URL through the security layer so crawlers can read it.","effort":"config","endpoint":null},"share_of_scans":{"pct":2.4,"scans":74,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_BLOCKED","api":"https://crawlcheck.io/api/rules?code=SITEMAP_BLOCKED","glossary":["https://crawlcheck.io/glossary/sitemap"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_BLOCKED"}},{"code":"SITEMAP_ORIGIN_ERROR","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"The origin returned a server error on the sitemap path","meaning":null,"severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The sitemap URL returned a server error. Regenerate it (most CMSs have a setting) and check the URL loads in a private window.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.2,"scans":6,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_ORIGIN_ERROR","api":"https://crawlcheck.io/api/rules?code=SITEMAP_ORIGIN_ERROR","glossary":["https://crawlcheck.io/glossary/sitemap","https://crawlcheck.io/glossary/sitemap-origin-error"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_ORIGIN_ERROR"}},{"code":"HOMEPAGE_REFUSED","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"The homepage refused this client, so every homepage section is unmeasured","meaning":null,"severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The homepage was refused. Check bot-protection and WAF rules; a challenge page served with a 200 status reads as the site to a crawler.","effort":"config","endpoint":null},"share_of_scans":{"pct":2,"scans":64,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HOMEPAGE_REFUSED","api":"https://crawlcheck.io/api/rules?code=HOMEPAGE_REFUSED","glossary":["https://crawlcheck.io/glossary/homepage-refused","https://crawlcheck.io/glossary/refusal"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HOMEPAGE_REFUSED"}},{"code":"HOMEPAGE_IS_INTERSTITIAL","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"The homepage answered with an interstitial, not a page","meaning":null,"severity":{"level":"critical"},"scored":false,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Remove or lower the interstitial (age gate, consent wall, JS challenge) for crawlers, or make sure the real content is in the first response.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.2,"scans":5,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HOMEPAGE_IS_INTERSTITIAL","api":"https://crawlcheck.io/api/rules?code=HOMEPAGE_IS_INTERSTITIAL","glossary":["https://crawlcheck.io/glossary/202-interstitial"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HOMEPAGE_IS_INTERSTITIAL"}},{"code":"HOMEPAGE_REDIRECTS_OFF_HOST","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"The homepage sends crawlers to a different hostname","meaning":"A request for this domain's homepage answers with a redirect to a hostname outside its own www/apex pair - a regional storefront, a checkout host, a vendor platform. The scanner does not follow it, and neither does a crawler attribute what it finds there to this domain: content, schema and entity signals on the other host belong to the other host, and this domain reads as a site with no homepage. If the redirect is deliberate, the domain still needs a page of its own that a crawler can read.","severity":{"level":"medium"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Point the domain at the site it is meant to serve, or make the redirect a single permanent hop to the canonical host.","effort":"config","endpoint":null},"share_of_scans":{"pct":2,"scans":62,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#HOMEPAGE_REDIRECTS_OFF_HOST","api":"https://crawlcheck.io/api/rules?code=HOMEPAGE_REDIRECTS_OFF_HOST","glossary":["https://crawlcheck.io/glossary/off-host-redirect"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=HOMEPAGE_REDIRECTS_OFF_HOST"}},{"code":"SITEMAP_UNREADABLE","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"Whether a sitemap exists could not be determined — every path was refused","meaning":"A refusal is not an absence. The paths where a sitemap would live answered with a block, so this scan cannot say whether one exists — and it does not guess. Confirm from a second network, and check server logs grouped by source IP: if the block is aimed at this scanner's address, the finding is about our access; if named crawlers see the same refusal, the sitemap is unreachable to them too.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Fix the sitemap XML so it parses: one <urlset> or <sitemapindex>, valid <loc> entries, no HTML wrapper.","effort":"build","endpoint":null},"share_of_scans":{"pct":0.5,"scans":16,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_UNREADABLE","api":"https://crawlcheck.io/api/rules?code=SITEMAP_UNREADABLE","glossary":["https://crawlcheck.io/glossary/sitemap","https://crawlcheck.io/glossary/sitemap-index"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_UNREADABLE"}},{"code":"DECLARED_SITEMAP_REFUSED","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"The sitemap robots.txt declares refused this client","meaning":null,"severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The sitemap named in robots.txt refuses crawlers. Allow it through the security layer.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.5,"scans":17,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#DECLARED_SITEMAP_REFUSED","api":"https://crawlcheck.io/api/rules?code=DECLARED_SITEMAP_REFUSED","glossary":["https://crawlcheck.io/glossary/sitemap"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=DECLARED_SITEMAP_REFUSED"}},{"code":"SITEMAP_DECLARES_NOTHING","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"A sitemap answers 200 and lists nothing","meaning":"The file parses as a sitemap and contains no URLs. A crawler that follows the declaration fetches it on every visit and learns nothing, and Search Console reports 0 discovered URLs for it. The usual cause is a generator with an empty source — an entity-map, image or news sitemap whose feed returned nothing — or a rewrite rule left behind. Either populate it or stop declaring it.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Populate the empty sitemap or remove it from robots.txt and from the sitemap index. A file that lists nothing should not be declared.","effort":"build","endpoint":null},"share_of_scans":{"pct":1.3,"scans":40,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#SITEMAP_DECLARES_NOTHING","api":"https://crawlcheck.io/api/rules?code=SITEMAP_DECLARES_NOTHING","glossary":["https://crawlcheck.io/glossary/sitemap"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=SITEMAP_DECLARES_NOTHING"}},{"code":"ORIGIN_BLOCKED_BEHIND_CACHE","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"A security layer refused this file — source unconfirmed","meaning":"The cached copy is fine; behind it the origin refused this request. IMPORTANT LIMIT: a scanner runs from one IP, and application firewalls block by IP reputation — so this may mean the site refuses crawlers, or it may only mean the site refuses THIS scanner. The two are indistinguishable from outside. Confirm in the server access logs, grouped by source IP, before acting on it.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"The cached copy works but the origin refuses. Fix the origin block before the cache expires, or the next miss serves the block to everyone.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.2,"scans":7,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ORIGIN_BLOCKED_BEHIND_CACHE","api":"https://crawlcheck.io/api/rules?code=ORIGIN_BLOCKED_BEHIND_CACHE","glossary":["https://crawlcheck.io/glossary/edge-cache","https://crawlcheck.io/glossary/cache-buster"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ORIGIN_BLOCKED_BEHIND_CACHE"}},{"code":"DECLARED_SITEMAP_BROKEN","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt points at a sitemap that fails","meaning":"robots.txt advertises a sitemap URL that does not resolve. Crawlers that trust the declaration and fail get nothing.","severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"robots.txt names a sitemap that does not resolve. Fix the URL in robots.txt or publish the sitemap at the path it names.","effort":"build","endpoint":null},"share_of_scans":{"pct":1.3,"scans":42,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#DECLARED_SITEMAP_BROKEN","api":"https://crawlcheck.io/api/rules?code=DECLARED_SITEMAP_BROKEN","glossary":["https://crawlcheck.io/glossary/sitemap","https://crawlcheck.io/glossary/sitemap-index"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=DECLARED_SITEMAP_BROKEN"}},{"code":"NO_SITEMAP_FOUND","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"No XML sitemap found","meaning":"Without a sitemap, discovery depends entirely on internal linking. Orphaned and recently published pages may never be found.","severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Publish a sitemap at /sitemap.xml (most CMSs generate one) and name it in robots.txt.","effort":"build","endpoint":null},"share_of_scans":{"pct":8.7,"scans":272,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#NO_SITEMAP_FOUND","api":"https://crawlcheck.io/api/rules?code=NO_SITEMAP_FOUND","glossary":["https://crawlcheck.io/glossary/sitemap"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=NO_SITEMAP_FOUND"}},{"code":"ROBOTS_NOT_200","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt does not return 200","meaning":"The response code on robots.txt determines how crawlers treat the whole site. Anything other than 200 or a clean 404 is risky.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Serve robots.txt with status 200 as text/plain. A 404 or 410 means no file: under RFC 9309 a crawler may then fetch everything, so any rule you meant to publish is not in force. A 429 is read by Google as a server error, which pauses crawling: exempt robots.txt from rate limits.","effort":"config","endpoint":null},"share_of_scans":{"pct":3.9,"scans":123,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_NOT_200","api":"https://crawlcheck.io/api/rules?code=ROBOTS_NOT_200","glossary":["https://crawlcheck.io/glossary/robots-txt","https://crawlcheck.io/glossary/robots-unavailable"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_NOT_200"}},{"code":"ROBOTS_NO_SITEMAP","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt declares no sitemap","meaning":"Adding one Sitemap: line lets crawlers enumerate the site instead of guessing at it.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Add a Sitemap: line to robots.txt pointing at the full sitemap URL, so crawlers find it without guessing.","effort":"config","endpoint":"https://crawlcheck.io/api/fix/robots?domain={domain}"},"share_of_scans":{"pct":9,"scans":281,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_NO_SITEMAP","api":"https://crawlcheck.io/api/rules?code=ROBOTS_NO_SITEMAP","glossary":["https://crawlcheck.io/glossary/robots-txt"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_NO_SITEMAP"}},{"code":"AI_OPTOUT_SET","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"An AI opt-out signal is set","meaning":"A Content-Signal directive is telling AI systems not to use this site. CDNs inject this by default on some plans, so it is frequently set without the owner deciding to set it.","severity":{"level":"medium"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"This is a choice, not a defect. If you want answer engines to quote the site, remove the AI opt-out directives; if you do not, leave it.","effort":"config","endpoint":null},"share_of_scans":{"pct":5,"scans":156,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#AI_OPTOUT_SET","api":"https://crawlcheck.io/api/rules?code=AI_OPTOUT_SET","glossary":["https://crawlcheck.io/glossary/ai-opt-out","https://crawlcheck.io/glossary/machine-readable-opt-out"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=AI_OPTOUT_SET"}},{"code":"STAGING_LEAK","kind":"scan","family":{"id":"site","label":"Site, hosts and page"},"title":"A staging hostname is exposed in the page source","meaning":"A staging hostname appears in the page source. If that hostname is reachable and indexable it competes with the production site for the same content.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Remove staging or development hostnames, paths and comments from the production HTML, and block the staging site from crawlers.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.1,"scans":3,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#STAGING_LEAK","api":"https://crawlcheck.io/api/rules?code=STAGING_LEAK","glossary":["https://crawlcheck.io/glossary/staging-leak"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=STAGING_LEAK"}},{"code":"ROBOTS_CTYPE","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"robots.txt has an unexpected content type","meaning":"robots.txt should be served as text/plain. Some parsers are strict about this.","severity":{"level":"low"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Serve robots.txt as text/plain. Some validators refuse any other content type.","effort":"config","endpoint":null},"share_of_scans":{"pct":0.1,"scans":3,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ROBOTS_CTYPE","api":"https://crawlcheck.io/api/rules?code=ROBOTS_CTYPE","glossary":["https://crawlcheck.io/glossary/robots-txt","https://crawlcheck.io/glossary/robots-parse-error"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ROBOTS_CTYPE"}},{"code":"PLACEHOLDER_CONTENT","kind":"scan","family":{"id":"mostly_code","label":"Mostly code"},"title":"Unfinished placeholder text is live on this page","meaning":"Template text that was never replaced is being served to visitors and to crawlers. An answer engine cannot tell draft copy from real copy — it reads and may quote whatever is there. Placeholder in a title or meta description is worse than in body copy, because that is the text a search result shows and an assistant repeats.","severity":{"level":"high"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Replace the placeholder text (lorem ipsum, 'test content', template copy) with the real page, or unpublish the page until it is written.","effort":"content","endpoint":null},"share_of_scans":{"pct":0.2,"scans":5,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#PLACEHOLDER_CONTENT","api":"https://crawlcheck.io/api/rules?code=PLACEHOLDER_CONTENT","glossary":["https://crawlcheck.io/glossary/placeholder-content"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=PLACEHOLDER_CONTENT"}},{"code":"NO_LLMS_TXT","kind":"scan","family":{"id":"machine_files","label":"Machine files"},"title":"No llms.txt","meaning":"Not a defect. llms.txt is an emerging convention for telling AI systems what a site is and which pages matter.","severity":{"level":"info"},"scored":true,"measured_on":"every scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Publish /llms.txt: a short Markdown file naming the business, what it does, and the pages worth reading, served as text/plain or text/markdown.","effort":"content","endpoint":"https://crawlcheck.io/api/tool/llms?domain={domain}&format=txt"},"share_of_scans":{"pct":30.9,"scans":970,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#NO_LLMS_TXT","api":"https://crawlcheck.io/api/rules?code=NO_LLMS_TXT","glossary":["https://crawlcheck.io/glossary/llms-txt","https://crawlcheck.io/glossary/machine-layer"],"workflow":"https://crawlcheck.io/workflow/NO_LLMS_TXT","lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=NO_LLMS_TXT"}},{"code":"ENTITY_COLLISION","kind":"scan","family":{"id":"jsonld","label":"JSON-LD and entities"},"title":"Two different entities share the same identifying data","meaning":"Two entities on the page resolve to the same identifying data (e.g. name, phone, or place ID), so they cannot both be correct at once.","severity":{"level":"medium"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Merge the duplicate business nodes into one, keep a single @id, and remove any plugin that emits a second empty LocalBusiness record.","effort":"build","endpoint":"https://crawlcheck.io/api/fix/entitymap?domain={domain}"},"share_of_scans":{"pct":3.5,"scans":110,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ENTITY_COLLISION","api":"https://crawlcheck.io/api/rules?code=ENTITY_COLLISION","glossary":["https://crawlcheck.io/glossary/entity-collision"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ENTITY_COLLISION"}},{"code":"ENTITY_NO_COORDINATES","kind":"scan","family":{"id":"jsonld","label":"JSON-LD and entities"},"title":"An entity is missing geographic coordinates","meaning":"An entity is missing latitude/longitude coordinates, so it cannot be verified against a real-world location.","severity":{"level":"info"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Add latitude and longitude to the LocalBusiness node's geo property so an engine can place the business on a map without guessing.","effort":"content","endpoint":"https://crawlcheck.io/api/fix/entitymap?domain={domain}"},"share_of_scans":{"pct":5.7,"scans":178,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ENTITY_NO_COORDINATES","api":"https://crawlcheck.io/api/rules?code=ENTITY_NO_COORDINATES","glossary":["https://crawlcheck.io/glossary/geocoordinates","https://crawlcheck.io/glossary/geocoding","https://crawlcheck.io/glossary/postaladdress"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ENTITY_NO_COORDINATES"}},{"code":"GEO_COUNTRY_MISMATCH","kind":"scan","family":{"id":"jsonld","label":"JSON-LD and entities"},"title":"Published coordinates are not in the country the address names","meaning":"The latitude/longitude on the page falls outside the country its own address names. A range check cannot catch this: a missing minus sign is still a legal coordinate. Coordinates are read only by machines, so the error produces no visible symptom.","severity":{"level":"low"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":0.2,"scans":6,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#GEO_COUNTRY_MISMATCH","api":"https://crawlcheck.io/api/rules?code=GEO_COUNTRY_MISMATCH","glossary":["https://crawlcheck.io/glossary/geocoordinates","https://crawlcheck.io/glossary/geocoding","https://crawlcheck.io/glossary/postaladdress"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=GEO_COUNTRY_MISMATCH"}},{"code":"NAP_NAME_DRIFT","kind":"scan","family":{"id":"nap","label":"NAP and claims"},"title":"A directory prints a different business name than the site declares","meaning":"Two names on one phone number read as two entities, or as one entity with an unsettled name; an engine reconciling them may print either.","severity":{"level":"low"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":0.7,"scans":21,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#NAP_NAME_DRIFT","api":"https://crawlcheck.io/api/rules?code=NAP_NAME_DRIFT","glossary":["https://crawlcheck.io/glossary/nap","https://crawlcheck.io/glossary/nap-consistency","https://crawlcheck.io/glossary/name-drift"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=NAP_NAME_DRIFT"}},{"code":"NAP_UNDECLARED_LISTING","kind":"scan","family":{"id":"nap","label":"NAP and claims"},"title":"Off-site listings carry this phone number and the site does not declare them","meaning":"A listing the site does not declare is one an engine corroborates without you: whatever it prints becomes your name, hours and category in the answer.","severity":{"level":"info"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":4.9,"scans":153,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#NAP_UNDECLARED_LISTING","api":"https://crawlcheck.io/api/rules?code=NAP_UNDECLARED_LISTING","glossary":["https://crawlcheck.io/glossary/nap","https://crawlcheck.io/glossary/nap-consistency"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=NAP_UNDECLARED_LISTING"}},{"code":"CITED_WITH_WRONG_FACTS","kind":"answer_correlation","family":{"id":"answers","label":"Answer-engine citations"},"title":"An answer engine cites the site with the wrong phone or address","meaning":"The engine read the business facts from somewhere other than the site, usually a directory, and is repeating them.","severity":{"level":"medium"},"scored":false,"measured_on":"answer runs","revision":{"current":1,"revised_at":null},"fix":{"advice":"Fix the source the engine read (usually a directory listing), then make the site's schema and sameAs agree with it.","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"measured in answer runs, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CITED_WITH_WRONG_FACTS","api":"https://crawlcheck.io/api/rules?code=CITED_WITH_WRONG_FACTS","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"RENDER_RESOURCE_BLOCKED","kind":"scan","family":{"id":"robots","label":"Robots"},"title":"The homepage depends on files this site's own robots.txt disallows","meaning":"A crawler that obeys robots.txt cannot load those files, so it renders a different page from the one visitors see.","severity":{"level":"medium"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":2,"revised_at":"2026-09-30"},"fix":{"advice":"Allow crawlers to fetch the CSS and JavaScript the page depends on. Remove the Disallow rules covering those paths, or serve the content without them.","effort":"config","endpoint":null},"share_of_scans":{"pct":2.6,"scans":83,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#RENDER_RESOURCE_BLOCKED","api":"https://crawlcheck.io/api/rules?code=RENDER_RESOURCE_BLOCKED","glossary":["https://crawlcheck.io/glossary/blocked-render-resource"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=RENDER_RESOURCE_BLOCKED"}},{"code":"REACHABLE_NOT_CITED","kind":"answer_correlation","family":{"id":"answers","label":"Answer-engine citations"},"title":"An answer engine can reach the site and does not cite it","meaning":"Access is not the problem: the content and authority are.","severity":{"level":"info"},"scored":false,"measured_on":"answer runs","revision":{"current":1,"revised_at":null},"fix":{"advice":"See the Quote pillar rows on the machine-layer report: defining lead, quotable paragraphs, FAQ parity.","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"measured in answer runs, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#REACHABLE_NOT_CITED","api":"https://crawlcheck.io/api/rules?code=REACHABLE_NOT_CITED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"BOT_BLOCKED_NOT_CITED","kind":"answer_correlation","family":{"id":"answers","label":"Answer-engine citations"},"title":"An answer engine does not cite the site and robots.txt blocks its crawler","meaning":"An engine cannot quote a page its crawler is told not to read.","severity":{"level":"medium"},"scored":false,"measured_on":"answer runs","revision":{"current":1,"revised_at":null},"fix":{"advice":"Allow the engine's crawler in robots.txt (User-agent: <crawler> / Allow: /).","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"measured in answer runs, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#BOT_BLOCKED_NOT_CITED","api":"https://crawlcheck.io/api/rules?code=BOT_BLOCKED_NOT_CITED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"BOT_RULES_SHADOWED_NOT_CITED","kind":"answer_correlation","family":{"id":"answers","label":"Answer-engine citations"},"title":"An answer engine does not cite the site and its crawler group is shadowed","meaning":"The crawler's own group in robots.txt does not carry the rules the site meant it to follow.","severity":{"level":"low"},"scored":false,"measured_on":"answer runs","revision":{"current":1,"revised_at":null},"fix":{"advice":"Move the named crawler group above the wildcard group, or repeat its Allow lines after it.","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"measured in answer runs, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#BOT_RULES_SHADOWED_NOT_CITED","api":"https://crawlcheck.io/api/rules?code=BOT_RULES_SHADOWED_NOT_CITED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"NO_LLMS_TXT_NOT_CITED","kind":"answer_correlation","family":{"id":"answers","label":"Answer-engine citations"},"title":"An answer engine does not cite the site and there is no llms.txt","meaning":"The engine can reach the site; the site gives it no summary to start from.","severity":{"level":"info"},"scored":false,"measured_on":"answer runs","revision":{"current":1,"revised_at":null},"fix":{"advice":"Publish /llms.txt (the scanner's /api/tool/llms drafts one from the sitemap).","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"measured in answer runs, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#NO_LLMS_TXT_NOT_CITED","api":"https://crawlcheck.io/api/rules?code=NO_LLMS_TXT_NOT_CITED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_DECLARED_PUBLIC_OBSERVED_AUTH","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability declared public observed auth","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_DECLARED_PUBLIC_OBSERVED_AUTH","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_DECLARED_PUBLIC_OBSERVED_AUTH","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_DECLARED_ENDPOINT_MISSING","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability declared endpoint missing","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_DECLARED_ENDPOINT_MISSING","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_DECLARED_ENDPOINT_MISSING","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_DECLARED_NO_REQUIRED_INPUT_OBSERVED_REJECTED","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability declared no required input observed rejected","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_DECLARED_NO_REQUIRED_INPUT_OBSERVED_REJECTED","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_DECLARED_NO_REQUIRED_INPUT_OBSERVED_REJECTED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_DECLARED_MEDIA_MISMATCH","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability declared media mismatch","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_DECLARED_MEDIA_MISMATCH","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_DECLARED_MEDIA_MISMATCH","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_DECLARED_TOOL_NOT_LISTED","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability declared tool not listed","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_DECLARED_TOOL_NOT_LISTED","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_DECLARED_TOOL_NOT_LISTED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"CAPABILITY_LISTED_TOOL_NOT_DECLARED","kind":"capability_mismatch","family":{"id":"capabilities","label":"Capabilities"},"title":"Declared capability listed tool not declared","meaning":null,"severity":{"level":null},"scored":false,"measured_on":"scans of sites that declare OpenAPI or MCP capabilities","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"listed in the record's capabilities section, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#CAPABILITY_LISTED_TOOL_NOT_DECLARED","api":"https://crawlcheck.io/api/rules?code=CAPABILITY_LISTED_TOOL_NOT_DECLARED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"FETCHER_DISAGREEMENT","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"Two of our fetch layers disagree about a machine file","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"medium"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#FETCHER_DISAGREEMENT","api":"https://crawlcheck.io/api/rules?code=FETCHER_DISAGREEMENT","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"BENCH_POOL_MISMATCH","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"The benchmark's pool size contradicts its scope","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"medium"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#BENCH_POOL_MISMATCH","api":"https://crawlcheck.io/api/rules?code=BENCH_POOL_MISMATCH","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"COHORT_BELOW_THRESHOLD","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"Ranked against a cohort under the minimum","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"low"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#COHORT_BELOW_THRESHOLD","api":"https://crawlcheck.io/api/rules?code=COHORT_BELOW_THRESHOLD","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"SECTION_SCORE_OUT_OF_RANGE","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"A section score outside 0 to 100","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"high"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#SECTION_SCORE_OUT_OF_RANGE","api":"https://crawlcheck.io/api/rules?code=SECTION_SCORE_OUT_OF_RANGE","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"RANKED_VALUE_NOT_REPRODUCIBLE","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"The ranked number does not follow from the printed table","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"high"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#RANKED_VALUE_NOT_REPRODUCIBLE","api":"https://crawlcheck.io/api/rules?code=RANKED_VALUE_NOT_REPRODUCIBLE","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"SECTION_WEIGHT_UNDECLARED","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"A section scored without a declared weight","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"medium"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#SECTION_WEIGHT_UNDECLARED","api":"https://crawlcheck.io/api/rules?code=SECTION_WEIGHT_UNDECLARED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"PATH_NOT_MEASURED","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"A deep path was asked for, the home page was measured","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"low"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#PATH_NOT_MEASURED","api":"https://crawlcheck.io/api/rules?code=PATH_NOT_MEASURED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"PLACEHOLDER_IN_RECORD","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"One of our placeholders reached the record","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"high"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#PLACEHOLDER_IN_RECORD","api":"https://crawlcheck.io/api/rules?code=PLACEHOLDER_IN_RECORD","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"GRADE_OFF_ITS_OWN_SCALE","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"The letter does not follow the record's own thresholds","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"high"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#GRADE_OFF_ITS_OWN_SCALE","api":"https://crawlcheck.io/api/rules?code=GRADE_OFF_ITS_OWN_SCALE","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"GRADED_THROUGH_A_WALL","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"A grade on a site that refused the scanner","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"critical"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":"The reading was taken through a block page. Allow the scanner and re-scan for a real grade.","effort":"config","endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#GRADED_THROUGH_A_WALL","api":"https://crawlcheck.io/api/rules?code=GRADED_THROUGH_A_WALL","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"EXTRACTORS_READ_NOTHING","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"Our extractors read nothing from a head that has tags","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"medium"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#EXTRACTORS_READ_NOTHING","api":"https://crawlcheck.io/api/rules?code=EXTRACTORS_READ_NOTHING","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"GRADE_MISSING","kind":"self_audit","family":{"id":"self_audit","label":"Self-audit: defects in our own record"},"title":"A graded record with no grade","meaning":"A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.","severity":{"level":"medium"},"scored":false,"measured_on":"every scan record, checked against itself before it is stored, and again by the verifier","revision":{"current":1,"revised_at":null},"fix":{"advice":null,"effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":3137,"basis":"about our own record, not counted per scan"},"links":{"self":"https://crawlcheck.io/rules#GRADE_MISSING","api":"https://crawlcheck.io/api/rules?code=GRADE_MISSING","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"MCP_TOOL_POISONING","kind":"mcp_server","family":{"id":"mcp","label":"MCP servers"},"title":"An MCP server's tool descriptions carry text aimed at the model","meaning":"Text a client hands the model from tools/list hides characters a person cannot see, tells the model to ignore its instructions or keep something from the user, asks for secrets or the conversation, or tells the model how to use other tools. A model reads it as an instruction; the person approving the server usually never sees it.","severity":{"level":"high"},"scored":false,"measured_on":"remote MCP servers in the official MCP Registry (tools/list read during the handshake; no tool is called), and any server or tool list sent to /api/v1/mcp/scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"Remove the text from the tool name, description and schema; descriptions should say what the tool does and nothing to the model about other tools, secrets or the user. Re-check with /api/v1/mcp/scan.","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":null,"basis":"measured on remote MCP servers in the official MCP Registry, not per site scan: see https://crawlcheck.io/api/mcp/index (tool_poisoning)"},"links":{"self":"https://crawlcheck.io/rules#MCP_TOOL_POISONING","api":"https://crawlcheck.io/api/rules?code=MCP_TOOL_POISONING","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}},{"code":"HIGH_RISK_FIELD_CHANGED","kind":"mcp_server","family":{"id":"watch","label":"High-risk field watch"},"title":"A high-risk field changed and the owner has not confirmed it","meaning":"One of the fields an attacker changes first moved since CrawlCheck last read it: the payment processors loaded on the homepage or checkout page, the host checkout links go to, a cryptocurrency wallet address printed on the page, the OAuth authorization server the site's MCP endpoint names, the domain's MX hosts, or its registry-listed MCP endpoints. Watched fields: payment_processors, checkout_host, wallet_addresses, oauth_authorization_server, mx_hosts, mcp_endpoints. The first reading is the baseline and never alerts.","severity":{"level":"high"},"scored":false,"measured_on":"remote MCP servers in the official MCP Registry (tools/list read during the handshake; no tool is called), and any server or tool list sent to /api/v1/mcp/scan","revision":{"current":1,"revised_at":null},"fix":{"advice":"If the change is yours, confirm it: POST /api/v1/high-risk/confirm {domain} with the key that holds the domain claim. If it is not yours, treat the site as compromised: flip the incident switch (POST /api/v1/incident) and restore the field.","effort":null,"endpoint":null},"share_of_scans":{"pct":null,"scans":null,"of_scans":null,"basis":"measured by the daily high-risk field watch, not per site scan: see https://crawlcheck.io/api/v1/high-risk"},"links":{"self":"https://crawlcheck.io/rules#HIGH_RISK_FIELD_CHANGED","api":"https://crawlcheck.io/api/rules?code=HIGH_RISK_FIELD_CHANGED","glossary":[],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":null}}]}