CrawlCheck

Glossary · Crawlers and access

homepage refused

A homepage answering 401, 403, 429 or similar to this client while robots.txt is served normally. The machine layer stays scored; every section that needs the homepage is unmeasured rather than scored, and the report says why. It is marked unconfirmed until checked from a second address, because a datacentre refusal may be about the scanner rather than about crawlers. Reported as HOMEPAGE_REFUSED.

Terms this definition uses

robots.txt · machine layer · refusal

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · FCrDNS · Crawler trap · Conditional request · ASN blocking · operator feed · residential proxy · Google-Extended · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · CCBot · Bytespider · agentic traffic · crawler classification

off-host redirect  ·  FCrDNS

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.