Glossary · Crawlers and access
crawler classification
The model edge providers adopted in 2026, Cloudflare from July with new-domain defaults changing on 15 September: verification confirms only who a crawler is, and access is then decided per behaviour class, Search, Agent, or Training, rather than per operator. A multi-purpose crawler is judged under every class it belongs to, which is how a rule that blocks Training can block a search crawler that also trains. It is the industry name for the identity-versus-permission distinction that verified crawler already implies. The classes are the provider's judgment of a crawler's purpose, and that judgment is not published in a form a site can audit.
Terms this definition uses
Crawlers and access
Who is fetching, whether they are who they claim, and what your rules actually permit.
retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · homepage refused · FCrDNS · Crawler trap · Conditional request · ASN blocking · operator feed · residential proxy · Google-Extended · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · CCBot · Bytespider · agentic traffic
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.