CrawlCheck

Methodology

The rulebook

Every finding the scanner can raise: what it means, how serious it is, how to fix it and how many scans carried it. Open a rule for the detail.

77rules
48affect the grade
7revised

/api/rules (JSON) · schema · compatibility policy · trust

Each finding has a level from info to critical. Open findings at medium and above lower the letter grade; fix them first. A rule that changes what it reports gets a new revision; a record keeps the revision that decided it, and a replay uses that revision, never today's.

Every rule

Site, hosts and page 8 Headers 4 Mostly code 3 Machine files 11 Robots 13 Delivery parity 8 JSON-LD and entities 3 NAP and claims 2 Answer-engine citations 5 Capabilities 6 Self-audit: defects in our own record 12 MCP servers 1 High-risk field watch 1

Site, hosts and page

CSP_REPORTING_BROKENCSP violation reporting is misconfiguredinfo0.4%rev 1
Why it matters
The Content-Security-Policy header declares where to send violation reports, and the declaration is wrong. `report-to` takes the NAME of a group defined in a separate header — given a URL it is silently ignored, so nothing is reported at all. And `report-uri` aimed at your own origin turns every violation into another request to the page you are already trying to keep fast. Neither breaks the page; both mean the reporting you think you have is either absent or self-inflicted.
Fix
The Content-Security-Policy report endpoint does not accept reports. Fix the report-uri / report-to target or remove the directive.
effort: config
Share of scans
0.4% — 11 of 3137 counted scans
Links
JSON · glossary: content-security-policy
HOST_DUPLICATE_200www and the apex both serve the site instead of one redirectingmedium3.5%rev 1
Why it matters
Both www and the bare domain answer 200 with the same pages, so every URL on this site exists at two addresses and a crawler fetches the whole site twice. A canonical tag does not prevent this: it is read after the response arrives, so the second fetch has already been spent. Nothing looks broken from a browser, which is why this survives for years. The fix is a 301 from one host to the other so there is one address per page.
Fix
Pick one host (apex or www) and redirect the other to it with a 301 so there is one copy of every page.
effort: config
Share of scans
3.5% — 111 of 3137 counted scans
Links
JSON · glossary: canonical · glossary: canonical-url · glossary: host-duplicate
ORIGIN_REFUSED_SCANNERThe homepage and robots.txt both refused this client, so the site was never measuredcritical0.3%rev 1
Fix
Allow this scanner's address through the security layer, or scan from the browser extension, which uses your own address.
effort: config
Share of scans
0.3% — 8 of 3137 counted scans
Links
JSON · glossary: refusal · glossary: datacenter-ip · glossary: vantage-point
DOMAIN_IS_ALIASThis domain is an alias: every path redirects to another sitelow0.7%rev 1
Why it matters
Every request to this hostname is redirected to a different site, so nothing measured here describes a site at this address. Scan the destination instead; the redirect itself is the only fact about this domain.
Fix
Nothing to fix on this domain: it forwards to another site. Scan the destination host instead.
Share of scans
0.7% — 21 of 3137 counted scans
Links
JSON · glossary: canonical · glossary: canonical-url · glossary: alias-domain
HOMEPAGE_REDIRECTS_OFF_HOSTThe homepage sends crawlers to a different hostnamemedium2%rev 1
Why it matters
A request for this domain's homepage answers with a redirect to a hostname outside its own www/apex pair - a regional storefront, a checkout host, a vendor platform. The scanner does not follow it, and neither does a crawler attribute what it finds there to this domain: content, schema and entity signals on the other host belong to the other host, and this domain reads as a site with no homepage. If the redirect is deliberate, the domain still needs a page of its own that a crawler can read.
Fix
Point the domain at the site it is meant to serve, or make the redirect a single permanent hop to the canonical host.
effort: config
Share of scans
2% — 62 of 3137 counted scans
Links
JSON · glossary: off-host-redirect
ORIGIN_BLOCKED_BEHIND_CACHEA security layer refused this file — source unconfirmedmedium0.2%rev 1
Why it matters
The cached copy is fine; behind it the origin refused this request. IMPORTANT LIMIT: a scanner runs from one IP, and application firewalls block by IP reputation — so this may mean the site refuses crawlers, or it may only mean the site refuses THIS scanner. The two are indistinguishable from outside. Confirm in the server access logs, grouped by source IP, before acting on it.
Fix
The cached copy works but the origin refuses. Fix the origin block before the cache expires, or the next miss serves the block to everyone.
effort: config
Share of scans
0.2% — 7 of 3137 counted scans
Links
JSON · glossary: edge-cache · glossary: cache-buster
STAGING_LEAKA staging hostname is exposed in the page sourcelow0.1%rev 1
Why it matters
A staging hostname appears in the page source. If that hostname is reachable and indexable it competes with the production site for the same content.
Fix
Remove staging or development hostnames, paths and comments from the production HTML, and block the staging site from crawlers.
effort: config
Share of scans
0.1% — 3 of 3137 counted scans
Links
JSON · glossary: staging-leak

Headers

HSTS_MISSINGHTTPS is served without a Strict-Transport-Security headerinfo9.3%rev 1
Why it matters
Strict-Transport-Security tells a browser that has visited once to use HTTPS for every later request without trying HTTP first. Without it, a typed or linked http:// address makes one plaintext request before the redirect. It does not change what crawlers receive or how the site ranks; it is a small, reversible hardening step, and it is listed at the lowest severity for that reason.
Fix
Send Strict-Transport-Security: max-age=15552000; includeSubDomains on HTTPS responses (Cloudflare: SSL/TLS, Edge Certificates, HSTS). Leave preload off unless you intend a change that takes months to undo.
generate the fix
effort: config
Share of scans
9.3% — 292 of 3137 counted scans
Links
JSON · glossary: hsts
HSTS_PRELOAD_INELIGIBLEThe HSTS header asks for browser preloading but does not qualify for itlow4.1%rev 1
Why it matters
The header carries the `preload` token, which is a request to be compiled into the HTTPS-only list shipped inside browsers. That list only accepts a header with max-age of at least 31536000 (one year) AND includeSubDomains AND preload. As sent, the token is ignored: the site is declaring an intent it does not meet, and if it ever was on the list, this header is what gets it removed. The rest of the HSTS header still works.
Fix
Either raise max-age to 31536000 and keep includeSubDomains so the preload token qualifies, or remove the preload token. As sent it does nothing.
generate the fix
effort: config
Share of scans
4.1% — 130 of 3137 counted scans
Links
JSON · glossary: hsts · glossary: hsts-preload-list
HSTS_PRELOAD_REMOVEDThis domain was removed from the browser HSTS preload listlow0.4%rev 1
Why it matters
The browser preload list (hstspreload.org) records this domain as removed. Domains are removed when the header stops meeting the requirements, or on request. Browsers that shipped with the old list still force HTTPS on every subdomain; newer ones do not, so behaviour now differs by browser version. Either the removal was intended and the `preload` token should go, or it was not and the header needs to qualify again before resubmitting.
Fix
If preloading was never intended, remove the preload token. If it was, restore a qualifying header (max-age at least 31536000, includeSubDomains, preload) and resubmit at hstspreload.org.
effort: config
Share of scans
0.4% — 14 of 3137 counted scans
Links
JSON · glossary: hsts · glossary: hsts-preload-list
STALE_CACHE_SERVEDVisitors and crawlers are being served an old copy of this pagelow17.5%rev 2 · 2026-09-30
Why it matters
The page handed to us carried a cache age older than an hour. The detail says whether the copy still matches what the origin produces now: if it matches, nothing is broken today; if it differs, an edit is already waiting behind the cache. It matters the moment you change something: an edit, a new price, a corrected phone number stays invisible to every visitor and every crawler for about that long, and nothing in your dashboard says so.
Fix
Lower the cache lifetime on HTML (an hour or less), or purge the edge and origin caches after every publish so crawlers see the current copy.
generate the fix
effort: config
Share of scans
17.5% — 549 of 3137 counted scans
Links
JSON · glossary: stale-response · glossary: edge-cache · glossary: age-header · the fix end to end on a real site

Mostly code

CONTENT_NEEDS_JAVASCRIPTThe page’s content only exists after JavaScript runsmedium1.1%rev 1
Why it matters
The HTML the server sends is an application shell: a framework mount point and a script bundle, with almost no readable text. A browser runs the bundle and the page appears. A crawler that does not execute JavaScript — which is most AI crawlers and every training-data fetch — receives the shell and quotes nothing. This is measured on the bytes the server delivered, not on a rendered view; a framework by itself is not the defect, an empty delivery is.
Fix
Render the page on the server or at build time (Next.js/Nuxt/SvelteKit SSR or SSG, or a prerender step) so the delivered HTML already contains the text. Check by fetching the page with curl: the words a crawler needs must be in that response.
effort: build
Share of scans
1.1% — 34 of 3137 counted scans
Links
JSON · glossary: render-dependence · glossary: render-dependence-gap · glossary: render-gap
PAGE_IS_MOSTLY_CODEAlmost nothing a machine receives from this page is readable textmedium29.2%rev 1
Why it matters
A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.
Fix
Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.
generate the fix
effort: build
Share of scans
29.2% — 916 of 3137 counted scans
Links
JSON · glossary: payload-text-share · the fix end to end on a real site
PLACEHOLDER_CONTENTUnfinished placeholder text is live on this pagehigh0.2%rev 1
Why it matters
Template text that was never replaced is being served to visitors and to crawlers. An answer engine cannot tell draft copy from real copy — it reads and may quote whatever is there. Placeholder in a title or meta description is worse than in body copy, because that is the text a search result shows and an assistant repeats.
Fix
Replace the placeholder text (lorem ipsum, 'test content', template copy) with the real page, or unpublish the page until it is written.
effort: content
Share of scans
0.2% — 5 of 3137 counted scans
Links
JSON · glossary: placeholder-content

Machine files

SITEMAP_NO_LASTMODThe sitemap does not say when anything changedlow6.6%rev 1
Why it matters
The sitemap lists URLs but carries no lastmod dates, so it says what exists and not what changed. A crawler with a budget uses lastmod to decide what to re-fetch; without it, the choice is made for you — usually by re-crawling little and late. This is the change signal every crawler already understands, and unlike a push protocol it needs no key, no account and no participation from the search engine.
Fix
Add <lastmod> dates to sitemap entries so a crawler can prioritise pages that changed.
generate the fix
effort: content
Share of scans
6.6% — 206 of 3137 counted scans
Links
JSON · glossary: sitemap · glossary: sitemap-lastmod
MACHINE_CHAIN_HEAVYThe machine files cost more than the pageinfo17.8%rev 2 · 2026-09-30
Why it matters
A crawler reads robots.txt, the sitemap, llms.txt and the agent files before the first page, and pays for every 404 that answers with a full page. Keep machine files as small as their job allows and serve absent paths with a short 404 body.
Fix
generate the fix
Share of scans
17.8% — 559 of 3137 counted scans
Links
JSON · glossary: machine-layer
MACHINE_FILE_CACHE_SPLITThis file answers differently depending on whether the cache is hithigh0.1%rev 1
Fix
Make the apex and www hosts serve the same machine files, or redirect one host to the other so there is only one copy to keep current.
effort: config
Share of scans
0.1% — 2 of 3137 counted scans
Links
JSON · glossary: edge-cache-pinning · glossary: cache-split · glossary: cache-buster
MACHINE_FILE_CACHE_STALEThe cache is still serving a file the origin no longer haslow0.6%rev 1
Fix
Give robots.txt, sitemap.xml and llms.txt a short cache lifetime (five minutes is plenty) so an edit reaches crawlers the same day.
effort: config
Share of scans
0.6% — 18 of 3137 counted scans
Links
JSON · glossary: cache-split · glossary: edge-cache · glossary: cache-buster
SITEMAP_IS_HTMLThe sitemap returns HTML, not XMLcritical0.7%rev 1
Why it matters
A crawler that cannot parse the sitemap cannot enumerate the site. It falls back to following links, so deep and newly published pages go undiscovered.
Fix
The sitemap URL is returning an HTML page. Serve XML at that path, or update robots.txt to point at the sitemap that actually exists.
effort: config
Share of scans
0.7% — 23 of 3137 counted scans
Links
JSON · glossary: sitemap · glossary: soft-404-machine-file
SITEMAP_BLOCKEDThe sitemap is refusedhigh2.4%rev 1
Why it matters
SAME LIMIT AS ABOVE: if a named application firewall appears in the response, this may be blocking the scanner's IP rather than crawlers. Confirm in server logs grouped by source IP. Otherwise: the server refuses this sitemap rather than returning it or a clean 404. Crawlers cannot enumerate the site and fall back to link-following, so deep and newly published pages go undiscovered.
Fix
Allow the sitemap URL through the security layer so crawlers can read it.
effort: config
Share of scans
2.4% — 74 of 3137 counted scans
Links
JSON · glossary: sitemap
SITEMAP_ORIGIN_ERRORThe origin returned a server error on the sitemap pathlow0.2%rev 1
Fix
The sitemap URL returned a server error. Regenerate it (most CMSs have a setting) and check the URL loads in a private window.
effort: config
Share of scans
0.2% — 6 of 3137 counted scans
Links
JSON · glossary: sitemap · glossary: sitemap-origin-error
SITEMAP_UNREADABLEWhether a sitemap exists could not be determined — every path was refusedlow0.5%rev 1
Why it matters
A refusal is not an absence. The paths where a sitemap would live answered with a block, so this scan cannot say whether one exists — and it does not guess. Confirm from a second network, and check server logs grouped by source IP: if the block is aimed at this scanner's address, the finding is about our access; if named crawlers see the same refusal, the sitemap is unreachable to them too.
Fix
Fix the sitemap XML so it parses: one <urlset> or <sitemapindex>, valid <loc> entries, no HTML wrapper.
effort: build
Share of scans
0.5% — 16 of 3137 counted scans
Links
JSON · glossary: sitemap · glossary: sitemap-index
SITEMAP_DECLARES_NOTHINGA sitemap answers 200 and lists nothinglow1.3%rev 1
Why it matters
The file parses as a sitemap and contains no URLs. A crawler that follows the declaration fetches it on every visit and learns nothing, and Search Console reports 0 discovered URLs for it. The usual cause is a generator with an empty source — an entity-map, image or news sitemap whose feed returned nothing — or a rewrite rule left behind. Either populate it or stop declaring it.
Fix
Populate the empty sitemap or remove it from robots.txt and from the sitemap index. A file that lists nothing should not be declared.
effort: build
Share of scans
1.3% — 40 of 3137 counted scans
Links
JSON · glossary: sitemap
NO_SITEMAP_FOUNDNo XML sitemap foundhigh8.7%rev 1
Why it matters
Without a sitemap, discovery depends entirely on internal linking. Orphaned and recently published pages may never be found.
Fix
Publish a sitemap at /sitemap.xml (most CMSs generate one) and name it in robots.txt.
effort: build
Share of scans
8.7% — 272 of 3137 counted scans
Links
JSON · glossary: sitemap
NO_LLMS_TXTNo llms.txtinfo30.9%rev 1
Why it matters
Not a defect. llms.txt is an emerging convention for telling AI systems what a site is and which pages matter.
Fix
Publish /llms.txt: a short Markdown file naming the business, what it does, and the pages worth reading, served as text/plain or text/markdown.
generate the fix
effort: content
Share of scans
30.9% — 970 of 3137 counted scans
Links
JSON · glossary: llms-txt · glossary: machine-layer · the fix end to end on a real site

Robots

ROBOTS_RULES_SHADOWEDrobots.txt rules do not apply to named agentsmedium13.4%rev 2 · 2026-09-30
Why it matters
A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.
Fix
Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.
generate the fix
effort: config
Share of scans
13.4% — 419 of 3137 counted scans
Links
JSON · glossary: robots-txt · glossary: robots-group-merging
INSTRUCTION_CONFLICTThe instruction files contradict each otherlow0.3%rev 1
Why it matters
A crawler obeys only its own User-agent group. A Content-Signal or Disallow written under * says nothing to an agent that has its own group, and a group that both invites and forbids leaves the crawler to pick. Put the signal in every group it is meant for, and let the rules under it agree.
Fix
the finding in your report names what to change
Share of scans
0.3% — 8 of 3137 counted scans
Links
JSON · glossary: x-robots-tag
ROBOTS_IS_HTMLrobots.txt is HTML, not textcritical0.2%rev 1
Why it matters
robots.txt must be plain text. Served as HTML it cannot be parsed, so no crawl directives and no sitemap reference are read.
Fix
robots.txt is returning an HTML page. Serve a real text file at /robots.txt, and check that no catch-all route is answering instead.
effort: config
Share of scans
0.2% — 6 of 3137 counted scans
Links
JSON · glossary: robots-txt · glossary: robots-parse-error · glossary: soft-404-machine-file
ROBOTS_IS_CATCHALLrobots.txt answers with the catch-all page - there is no robots.txtlow0.6%rev 1
Why it matters
The site answers robots.txt with the same page it serves for a path that does not exist. Crawlers cannot parse HTML as robots rules, so they proceed as if no file existed: everything allowed. Publish a real text/plain robots.txt.
Fix
A catch-all route returns the homepage for any path, including robots.txt. Add a real robots.txt and return 404 for paths that do not exist.
effort: config
Share of scans
0.6% — 20 of 3137 counted scans
Links
JSON · glossary: robots-txt
ROBOTS_BLOCKEDrobots.txt is unreadablehigh0.5%rev 1
Why it matters
A 5xx on robots.txt tells a standards-following crawler to treat the whole site as disallowed (RFC 9309). A 401 or 403 is a 4xx, which the standard treats as no file and so allows everything, though some crawlers read it as a refusal. Either way the rules you wrote are not the rules a crawler read. This is the single most expensive file on the site to get wrong.
Fix
robots.txt could not be fetched. Allow it through the security layer; it is the one file every crawler reads first.
effort: config
Share of scans
0.5% — 15 of 3137 counted scans
Links
JSON · glossary: robots-txt · glossary: robots-unavailable
ROBOTS_DISALLOW_ALLrobots.txt blocks every crawler it does not namecritical1.5%rev 2 · 2026-09-30
Why it matters
robots.txt instructs every crawler to stay off the entire site.
Fix
robots.txt disallows everything. Remove the blanket Disallow: / rule (or restrict it to the paths you actually mean).
effort: config
Share of scans
1.5% — 48 of 3137 counted scans
Links
JSON · glossary: robots-txt
DECLARED_SITEMAP_REFUSEDThe sitemap robots.txt declares refused this clientlow0.5%rev 1
Fix
The sitemap named in robots.txt refuses crawlers. Allow it through the security layer.
effort: config
Share of scans
0.5% — 17 of 3137 counted scans
Links
JSON · glossary: sitemap
DECLARED_SITEMAP_BROKENrobots.txt points at a sitemap that failshigh1.3%rev 1
Why it matters
robots.txt advertises a sitemap URL that does not resolve. Crawlers that trust the declaration and fail get nothing.
Fix
robots.txt names a sitemap that does not resolve. Fix the URL in robots.txt or publish the sitemap at the path it names.
effort: build
Share of scans
1.3% — 42 of 3137 counted scans
Links
JSON · glossary: sitemap · glossary: sitemap-index
ROBOTS_NOT_200robots.txt does not return 200medium3.9%rev 1
Why it matters
The response code on robots.txt determines how crawlers treat the whole site. Anything other than 200 or a clean 404 is risky.
Fix
Serve robots.txt with status 200 as text/plain. A 404 or 410 means no file: under RFC 9309 a crawler may then fetch everything, so any rule you meant to publish is not in force. A 429 is read by Google as a server error, which pauses crawling: exempt robots.txt from rate limits.
effort: config
Share of scans
3.9% — 123 of 3137 counted scans
Links
JSON · glossary: robots-txt · glossary: robots-unavailable
ROBOTS_NO_SITEMAProbots.txt declares no sitemapmedium9%rev 1
Why it matters
Adding one Sitemap: line lets crawlers enumerate the site instead of guessing at it.
Fix
Add a Sitemap: line to robots.txt pointing at the full sitemap URL, so crawlers find it without guessing.
generate the fix
effort: config
Share of scans
9% — 281 of 3137 counted scans
Links
JSON · glossary: robots-txt
AI_OPTOUT_SETAn AI opt-out signal is setmedium5%rev 1
Why it matters
A Content-Signal directive is telling AI systems not to use this site. CDNs inject this by default on some plans, so it is frequently set without the owner deciding to set it.
Fix
This is a choice, not a defect. If you want answer engines to quote the site, remove the AI opt-out directives; if you do not, leave it.
effort: config
Share of scans
5% — 156 of 3137 counted scans
Links
JSON · glossary: ai-opt-out · glossary: machine-readable-opt-out
ROBOTS_CTYPErobots.txt has an unexpected content typelow0.1%rev 1
Why it matters
robots.txt should be served as text/plain. Some parsers are strict about this.
Fix
Serve robots.txt as text/plain. Some validators refuse any other content type.
effort: config
Share of scans
0.1% — 3 of 3137 counted scans
Links
JSON · glossary: robots-txt · glossary: robots-parse-error
RENDER_RESOURCE_BLOCKEDThe homepage depends on files this site's own robots.txt disallowsmedium2.6%rev 2 · 2026-09-30
Why it matters
A crawler that obeys robots.txt cannot load those files, so it renders a different page from the one visitors see.
Fix
Allow crawlers to fetch the CSS and JavaScript the page depends on. Remove the Disallow rules covering those paths, or serve the content without them.
effort: config
Share of scans
2.6% — 83 of 3137 counted scans
Links
JSON · glossary: blocked-render-resource

Delivery parity

CHALLENGE_SERVED_200A verification page is being served instead of the filecritical0.6%rev 1
Why it matters
The server answered with HTTP 200 — so uptime monitors, Search Console and your host all report this as healthy — but the body is a JavaScript verification page, not the file. Crawlers do not run JavaScript on these requests. They receive the interstitial and move on.
Fix
A verification page must not answer with status 200. Return 403 or 503 for challenges so crawlers do not index the challenge as the page.
effort: config
Share of scans
0.6% — 20 of 3137 counted scans
Links
JSON · glossary: challenge-page-at-200
ANSWER_ENGINE_REFUSEDAn answer engine's crawler was refused while a browser was servedhigh5.9%rev 1
Why it matters
The same URL, the same minute: an ordinary client was served the page and a named answer-engine crawler was not. That is the edge rather than a robots.txt decision. One caveat this site's own guides state and this finding must too: the scan sent the crawler's name from an address that operator does not publish, and an edge that verifies identity by IP is right to refuse it. So this proves the edge treats the identity differently - not that it refuses the real crawler. Your own logs, checked against the operator's published ranges, settle which. If the edge is refusing the real crawler, nothing else in this report reaches that engine, however good it is.
Fix
Allow the named answer-engine crawlers through the security layer (Cloudflare bot rules, WAF, rate limits), or verify their published IP ranges instead of blocking by name.
effort: config
Share of scans
5.9% — 184 of 3137 counted scans
Links
JSON · glossary: refusal · glossary: delivery-comparison
UNIFORM_REFUSALThis origin refused the scanner, so the site was never measuredcritical3.6%rev 1
Fix
The site refused every request from this scanner. Allow this scanner's address, or run the browser extension from your own network, then re-scan.
effort: config
Share of scans
3.6% — 112 of 3137 counted scans
Links
JSON · glossary: uniform-refusal · glossary: datacenter-ip
CRAWLER_SERVED_LESSA named crawler was served materially less text than an ordinary clienthigh0.4%rev 2 · 2026-09-30
Why it matters
Same URL, same minute, both answered 200 — but the named crawler received materially fewer words than an unnamed client did. A refusal is loud and easy to find; being served a thinner page is silent, and it is the version an answer engine quotes from.
Fix
Serve every crawler the same HTML a browser gets. Check bot-management rules, 'lite' templates and user-agent conditions in the CMS or CDN.
effort: config
Share of scans
0.4% — 14 of 3137 counted scans
Links
JSON · glossary: delivery-comparison
CRAWLER_REDIRECTED_AWAYA named crawler was sent somewhere the ordinary client was nothigh0.4%rev 2 · 2026-09-30
Why it matters
The ordinary client stayed on the URL and the named crawler was sent elsewhere. Whatever is at the destination is what that engine will read and quote instead of this page.
Fix
Remove the user-agent based redirect so named crawlers land on the page a visitor lands on.
effort: config
Share of scans
0.4% — 11 of 3137 counted scans
Links
JSON · glossary: delivery-comparison
CHALLENGE_PINNED_AT_EDGEThe verification page is cached at the CDN edgecritical0.5%rev 1
Why it matters
The origin marked this response private and uncacheable. A cache rule with a TTL override stored it anyway, so the verification page is now being handed to every visitor and every crawler from the edge until it expires or is purged.
Fix
Purge the edge cache: a challenge page was cached in place of the real file. Then exclude machine files from HTML caching rules.
effort: config
Share of scans
0.5% — 15 of 3137 counted scans
Links
JSON · glossary: edge-cache-pinning · glossary: challenge-page-at-200
HOMEPAGE_REFUSEDThe homepage refused this client, so every homepage section is unmeasuredlow2%rev 1
Fix
The homepage was refused. Check bot-protection and WAF rules; a challenge page served with a 200 status reads as the site to a crawler.
effort: config
Share of scans
2% — 64 of 3137 counted scans
Links
JSON · glossary: homepage-refused · glossary: refusal
HOMEPAGE_IS_INTERSTITIALThe homepage answered with an interstitial, not a pagecritical0.2%rev 1
Fix
Remove or lower the interstitial (age gate, consent wall, JS challenge) for crawlers, or make sure the real content is in the first response.
effort: config
Share of scans
0.2% — 5 of 3137 counted scans
Links
JSON · glossary: 202-interstitial

JSON-LD and entities

ENTITY_COLLISIONTwo different entities share the same identifying datamedium3.5%rev 1
Why it matters
Two entities on the page resolve to the same identifying data (e.g. name, phone, or place ID), so they cannot both be correct at once.
Fix
Merge the duplicate business nodes into one, keep a single @id, and remove any plugin that emits a second empty LocalBusiness record.
generate the fix
effort: build
Share of scans
3.5% — 110 of 3137 counted scans
Links
JSON · glossary: entity-collision
ENTITY_NO_COORDINATESAn entity is missing geographic coordinatesinfo5.7%rev 1
Why it matters
An entity is missing latitude/longitude coordinates, so it cannot be verified against a real-world location.
Fix
Add latitude and longitude to the LocalBusiness node's geo property so an engine can place the business on a map without guessing.
generate the fix
effort: content
Share of scans
5.7% — 178 of 3137 counted scans
Links
JSON · glossary: geocoordinates · glossary: geocoding · glossary: postaladdress
GEO_COUNTRY_MISMATCHPublished coordinates are not in the country the address nameslow0.2%rev 1
Why it matters
The latitude/longitude on the page falls outside the country its own address names. A range check cannot catch this: a missing minus sign is still a legal coordinate. Coordinates are read only by machines, so the error produces no visible symptom.
Fix
the finding in your report names what to change
Share of scans
0.2% — 6 of 3137 counted scans
Links
JSON · glossary: geocoordinates · glossary: geocoding · glossary: postaladdress

NAP and claims

NAP_NAME_DRIFTA directory prints a different business name than the site declareslow0.7%rev 1
Why it matters
Two names on one phone number read as two entities, or as one entity with an unsettled name; an engine reconciling them may print either.
Fix
the finding in your report names what to change
Share of scans
0.7% — 21 of 3137 counted scans
Links
JSON · glossary: nap · glossary: nap-consistency · glossary: name-drift
NAP_UNDECLARED_LISTINGOff-site listings carry this phone number and the site does not declare theminfo4.9%rev 1
Why it matters
A listing the site does not declare is one an engine corroborates without you: whatever it prints becomes your name, hours and category in the answer.
Fix
the finding in your report names what to change
Share of scans
4.9% — 153 of 3137 counted scans
Links
JSON · glossary: nap · glossary: nap-consistency

Answer-engine citations

CITED_WITH_WRONG_FACTSAn answer engine cites the site with the wrong phone or addressmedium—rev 1
Why it matters
The engine read the business facts from somewhere other than the site, usually a directory, and is repeating them.
Fix
Fix the source the engine read (usually a directory listing), then make the site's schema and sameAs agree with it.
Share of scans
measured in answer runs, not counted per scan
Links
JSON
REACHABLE_NOT_CITEDAn answer engine can reach the site and does not cite itinfo—rev 1
Why it matters
Access is not the problem: the content and authority are.
Fix
See the Quote pillar rows on the machine-layer report: defining lead, quotable paragraphs, FAQ parity.
Share of scans
measured in answer runs, not counted per scan
Links
JSON
BOT_BLOCKED_NOT_CITEDAn answer engine does not cite the site and robots.txt blocks its crawlermedium—rev 1
Why it matters
An engine cannot quote a page its crawler is told not to read.
Fix
Allow the engine's crawler in robots.txt (User-agent: <crawler> / Allow: /).
Share of scans
measured in answer runs, not counted per scan
Links
JSON
BOT_RULES_SHADOWED_NOT_CITEDAn answer engine does not cite the site and its crawler group is shadowedlow—rev 1
Why it matters
The crawler's own group in robots.txt does not carry the rules the site meant it to follow.
Fix
Move the named crawler group above the wildcard group, or repeat its Allow lines after it.
Share of scans
measured in answer runs, not counted per scan
Links
JSON
NO_LLMS_TXT_NOT_CITEDAn answer engine does not cite the site and there is no llms.txtinfo—rev 1
Why it matters
The engine can reach the site; the site gives it no summary to start from.
Fix
Publish /llms.txt (the scanner's /api/tool/llms drafts one from the sitemap).
Share of scans
measured in answer runs, not counted per scan
Links
JSON

Capabilities

CAPABILITY_DECLARED_PUBLIC_OBSERVED_AUTHDeclared capability declared public observed authnot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON
CAPABILITY_DECLARED_ENDPOINT_MISSINGDeclared capability declared endpoint missingnot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON
CAPABILITY_DECLARED_NO_REQUIRED_INPUT_OBSERVED_REJECTEDDeclared capability declared no required input observed rejectednot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON
CAPABILITY_DECLARED_MEDIA_MISMATCHDeclared capability declared media mismatchnot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON
CAPABILITY_DECLARED_TOOL_NOT_LISTEDDeclared capability declared tool not listednot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON
CAPABILITY_LISTED_TOOL_NOT_DECLAREDDeclared capability listed tool not declarednot scored—rev 1
Fix
the finding in your report names what to change
Share of scans
listed in the record's capabilities section, not counted per scan
Links
JSON

Self-audit: defects in our own record

FETCHER_DISAGREEMENTTwo of our fetch layers disagree about a machine filemedium—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
BENCH_POOL_MISMATCHThe benchmark's pool size contradicts its scopemedium—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
COHORT_BELOW_THRESHOLDRanked against a cohort under the minimumlow—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
SECTION_SCORE_OUT_OF_RANGEA section score outside 0 to 100high—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
RANKED_VALUE_NOT_REPRODUCIBLEThe ranked number does not follow from the printed tablehigh—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
SECTION_WEIGHT_UNDECLAREDA section scored without a declared weightmedium—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
PATH_NOT_MEASUREDA deep path was asked for, the home page was measuredlow—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
PLACEHOLDER_IN_RECORDOne of our placeholders reached the recordhigh—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
GRADE_OFF_ITS_OWN_SCALEThe letter does not follow the record's own thresholdshigh—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
GRADED_THROUGH_A_WALLA grade on a site that refused the scannercritical—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
The reading was taken through a block page. Allow the scanner and re-scan for a real grade.
effort: config
Share of scans
about our own record, not counted per scan
Links
JSON
EXTRACTORS_READ_NOTHINGOur extractors read nothing from a head that has tagsmedium—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON
GRADE_MISSINGA graded record with no grademedium—rev 1
Why it matters
A contradiction inside CrawlCheck's own record: a defect in the scanner, never in the scanned site.
Fix
the finding in your report names what to change
Share of scans
about our own record, not counted per scan
Links
JSON

MCP servers

MCP_TOOL_POISONINGAn MCP server's tool descriptions carry text aimed at the modelhigh—rev 1
Why it matters
Text a client hands the model from tools/list hides characters a person cannot see, tells the model to ignore its instructions or keep something from the user, asks for secrets or the conversation, or tells the model how to use other tools. A model reads it as an instruction; the person approving the server usually never sees it.
Fix
Remove the text from the tool name, description and schema; descriptions should say what the tool does and nothing to the model about other tools, secrets or the user. Re-check with /api/v1/mcp/scan.
Share of scans
measured on remote MCP servers in the official MCP Registry, not per site scan: see https://crawlcheck.io/api/mcp/index (tool_poisoning)
Links
JSON

High-risk field watch

HIGH_RISK_FIELD_CHANGEDA high-risk field changed and the owner has not confirmed ithigh—rev 1
Why it matters
One of the fields an attacker changes first moved since CrawlCheck last read it: the payment processors loaded on the homepage or checkout page, the host checkout links go to, a cryptocurrency wallet address printed on the page, the OAuth authorization server the site's MCP endpoint names, the domain's MX hosts, or its registry-listed MCP endpoints. Watched fields: payment_processors, checkout_host, wallet_addresses, oauth_authorization_server, mx_hosts, mcp_endpoints. The first reading is the baseline and never alerts.
Fix
If the change is yours, confirm it: POST /api/v1/high-risk/confirm {domain} with the key that holds the domain claim. If it is not yours, treat the site as compromised: flip the incident switch (POST /api/v1/incident) and restore the field.
Share of scans
measured by the daily high-risk field watch, not per site scan: see https://crawlcheck.io/api/v1/high-risk
Links
JSON

Generated 2026-10-07 16:48 UTC. A rule you think is wrong? Write to hello@crawlcheck.io with the report id; corrections are published as a new revision.