{"ok":true,"kind":"crawlcheck-rulebook","v":1,"generated_at":"2026-10-07T17:43:58.250Z","score_version":28,"schema":"https://crawlcheck.io/schemas/rulebook.json","counts":{"rules":77,"by_kind":{"scan":52,"answer_correlation":5,"capability_mismatch":6,"self_audit":12,"mcp_server":2},"by_family":{"site":8,"headers":4,"mostly_code":3,"machine_files":11,"robots":13,"delivery":8,"jsonld":3,"nap":2,"answers":5,"capabilities":6,"self_audit":12,"mcp":1,"watch":1},"revised":7,"scored":48},"scans_counted":3137,"severity_scale":[{"level":"info"},{"level":"low"},{"level":"medium"},{"level":"high"},{"level":"critical"}],"method":{"revisions":"A rule that changes what it reports gets a new revision; a record keeps the revision that decided it, and a replay uses that revision, never today's.","share":"Share of counted scans that carried the code at least once, the same numbers /data publishes. Self-scans and opted-out sites are never counted.","grade":"Each finding has a level from info to critical. Open findings at medium and above lower the letter grade; fix them first."},"rules":[{"code":"ANSWER_ENGINE_REFUSED","kind":"scan","family":{"id":"delivery","label":"Delivery parity"},"title":"An answer engine's crawler was refused while a browser was served","meaning":"The same URL, the same minute: an ordinary client was served the page and a named answer-engine crawler was not. That is the edge rather than a robots.txt decision. One caveat this site's own guides state and this finding must too: the scan sent the crawler's name from an address that operator does not publish, and an edge that verifies identity by IP is right to refuse it. So this proves the edge treats the identity differently - not that it refuses the real crawler. Your own logs, checked against the operator's published ranges, settle which. If the edge is refusing the real crawler, nothing else in this report reaches that engine, however good it is.","severity":{"level":"high"},"scored":true,"measured_on":"full scans only (the daily watch does not measure it)","revision":{"current":1,"revised_at":null},"fix":{"advice":"Allow the named answer-engine crawlers through the security layer (Cloudflare bot rules, WAF, rate limits), or verify their published IP ranges instead of blocking by name.","effort":"config","endpoint":null},"share_of_scans":{"pct":5.9,"scans":184,"of_scans":3137,"basis":"share of counted scans that carried it; the same numbers /data publishes"},"links":{"self":"https://crawlcheck.io/rules#ANSWER_ENGINE_REFUSED","api":"https://crawlcheck.io/api/rules?code=ANSWER_ENGINE_REFUSED","glossary":["https://crawlcheck.io/glossary/refusal","https://crawlcheck.io/glossary/delivery-comparison"],"workflow":null,"lineage":"https://crawlcheck.io/docs/lineage-coverage","explain_template":"https://crawlcheck.io/api/explain?id={report_id}&code=ANSWER_ENGINE_REFUSED"}}],"filter":{"code":"ANSWER_ENGINE_REFUSED"}}