CrawlCheck

Glossary · Crawlers and access

operator feed

The publisher-side list of source addresses a crawler operator maintains, usually an IP-range JSON, a reverse-DNS suffix, or both. It is the only authoritative basis for calling a hit verified or forged: a request from an address in the feed is the operator's, a request outside it is not, whatever the user-agent string says. Meta, Apple and Amazon publish only a documentation page, so their hits can be checked by reverse DNS but not by range. A stale feed indicts the operator's publishing, not the site being measured.

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · homepage refused · FCrDNS · Crawler trap · Conditional request · ASN blocking · residential proxy · Google-Extended · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · CCBot · Bytespider · agentic traffic · crawler classification

ASN blocking  ·  residential proxy

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.