Glossary · Crawlers and access
CCBot
The Common Crawl crawler. Its open corpus is downloaded and used by many model builders, so a page can enter training sets without ever being fetched by any model operator's own crawler. That is the case that breaks the sentence we blocked GPTBot, so we are out of the training data: the copy that trained the model may have been taken by a nonprofit archive years earlier. CCBot honours robots.txt and publishes its identity, so exclusion is possible; it is only retroactive removal that is not.
Terms this definition uses
Corpus · training data · robots.txt
Crawlers and access
Who is fetching, whether they are who they claim, and what your rules actually permit.
retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · homepage refused · FCrDNS · Crawler trap · Conditional request · ASN blocking · operator feed · residential proxy · Google-Extended · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · Bytespider · agentic traffic · crawler classification
← PerplexityBot vs Perplexity-User · Bytespider →
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.