CrawlCheck

Glossary · Crawlers and access

Bytespider

ByteDance's training crawler, reported over several years by publishers and by edge providers to fetch without regard to robots.txt exclusions. It is the standing example that a robots.txt rule is a request, honoured by operators who choose to honour it and enforced by nothing. The practical consequence for measurement is ordering: verify who actually fetched before attributing volume to a name, because a crawler that ignores rules is also the easiest name to forge.

Terms this definition uses

robots.txt

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · homepage refused · FCrDNS · Crawler trap · Conditional request · ASN blocking · operator feed · residential proxy · Google-Extended · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · CCBot · agentic traffic · crawler classification

CCBot  ·  agentic traffic

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.