CrawlCheck

Glossary · Crawlers and access

Google-Extended

The robots.txt token that controls whether Google may use a site's content to train and ground its Gemini models. It is not a crawler; Googlebot still fetches the page. Blocking it does not remove a site from Google Search, from Discover, or from the retrieval path behind AI Overviews and AI Mode, which run on the Search index. It is the most commonly misread token on the web: sites block it expecting to leave AI answers and instead leave only the training set.

Terms this definition uses

robots.txt · token · retrieval · AI Overview · AI Mode

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

retrieval crawler · user-triggered fetch · verified crawler · forged crawler identity · unverifiable · robots.txt · user-agent group · AI opt-out · Content-Signal · crawl budget · cloaking · challenge page at 200 · uniform refusal · nonexistent-path control · blocked render resource · off-host redirect · homepage refused · FCrDNS · Crawler trap · Conditional request · ASN blocking · operator feed · residential proxy · Google-Agent · GPTBot vs OAI-SearchBot · ClaudeBot vs Claude-User · PerplexityBot vs Perplexity-User · CCBot · Bytespider · agentic traffic · crawler classification

residential proxy  ·  Google-Agent

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.