Reference
AI crawlers: who sends them, what each one is for, and how to verify them
In one sentence
This is the reference list of the 114 named crawler agents CrawlCheck resolves from a site’s robots.txt and probes for, grouped by what each one is for — answer engines that live-fetch pages to quote them, search indexers, training-data collectors, regional search, social previews and SEO tools — with the operator, the user-agent string, and how to verify a request against the IP ranges the operator publishes.
Every crawler this scanner recognises — 114 policy identities resolved from robots.txt, of which 12 are actively sent to your homepage on every scan — grouped by what they are actually for. The same company usually sends three different crawlers for three different purposes, and blocking one does not do what people expect the others to do.
This page is the reference. The measurement is separate: a scan sends 15 client identities to your homepage from one address inside one second — GPTBot, ClaudeBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot, Amazonbot, Bytespider, Meta-ExternalAgent, CCBot, plus an unnamed client and two mobile-browser controls; the other named agents below are resolved from your robots.txt, not sent — and records what each one was actually served. Same second, same address, so a difference between two of them is your edge deciding, not the internet being busy. It also resolves every named agent in your own robots.txt into what it is permitted to do, which is how a policy someone wrote two years ago turns out to say something they did not intend.
Quick facts for every named crawler
Operator, purpose, and whether that operator publishes the IP ranges that make a claimed identity checkable. A crawler with no published feed is unverifiable here — never counted as forged.
- Crawler
GPTBot- Operator
- OpenAI
- What it is for
- Training and index crawl for ChatGPT
- Published IP ranges
- openai.com/gptbot.json
- Crawler
OAI-SearchBot- Operator
- OpenAI
- What it is for
- Search index behind ChatGPT answers
- Published IP ranges
- openai.com/searchbot.json
- Crawler
ChatGPT-User- Operator
- OpenAI
- What it is for
- Fetches a page because a user asked for it in a chat
- Published IP ranges
- openai.com/chatgpt-user.json
- Crawler
ClaudeBot- Operator
- Anthropic
- What it is for
- Index and training crawl
- Published IP ranges
- claude.com/crawling/bots.json
- Crawler
Claude-User- Operator
- Anthropic
- What it is for
- Fetches a page on a user's request
- Published IP ranges
- claude.com/crawling/bots.json
- Crawler
Claude-SearchBot- Operator
- Anthropic
- What it is for
- Search retrieval
- Published IP ranges
- claude.com/crawling/bots.json
- Crawler
PerplexityBot- Operator
- Perplexity
- What it is for
- Index crawl
- Published IP ranges
- www.perplexity.ai/perplexitybot.json
- Crawler
Perplexity-User- Operator
- Perplexity
- What it is for
- User-triggered fetch
- Published IP ranges
- www.perplexity.ai/perplexity-user.json
- Crawler
Googlebot- Operator
- What it is for
- Search index crawl
- Published IP ranges
- developers.google.com/static/crawling/ipranges/common-crawlers.json
- Crawler
Google-Extended- Operator
- What it is for
- Training permission token, not a separate crawler
- Published IP ranges
- robots.txt control token only — it never appears in a request, so there is no request to verify
- Crawler
Bingbot- Operator
- Microsoft
- What it is for
- Search index crawl
- Published IP ranges
- www.bing.com/toolbox/bingbot.json
- Crawler
Applebot- Operator
- Apple
- What it is for
- Search and Siri index crawl
- Published IP ranges
- search.developer.apple.com/applebot.json
- Crawler
Applebot-Extended- Operator
- Apple
- What it is for
- Training permission token, not a separate crawler
- Published IP ranges
- search.developer.apple.com/applebot.json
- Crawler
Amazonbot- Operator
- Amazon
- What it is for
- Alexa and shopping index crawl
- Published IP ranges
- developer.amazon.com/amazonbot/ip-addresses/
- Crawler
Meta-ExternalAgent- Operator
- Meta
- What it is for
- Training and preview fetch
- Published IP ranges
- no JSON feed — Meta publishes its crawler ranges as ASN AS32934 (whois -h whois.radb.net -- '-i origin AS32934')
- Crawler
Bytespider- Operator
- ByteDance
- What it is for
- Training crawl
- Published IP ranges
- none published — claims from this identity are unverifiable
- Crawler
CCBot- Operator
- Common Crawl
- What it is for
- Open crawl corpus many models train on
- Published IP ranges
- index.commoncrawl.org/ccbot.json
- Crawler
Amzn-SearchBot- Operator
- Amazon
- What it is for
- Amazon search eligibility (Alexa and shopping search); not used for generative-AI training
- Published IP ranges
- developer.amazon.com/amazonbot/searchbot-ip-addresses/
- Crawler
Amzn-User- Operator
- Amazon
- What it is for
- Live fetch for a user action such as an Alexa question
- Published IP ranges
- developer.amazon.com/amazonbot/live-ip-addresses/
- Crawler
MistralAI-Index- Operator
- Mistral
- What it is for
- Index crawl for Mistral search and Vibe answers; not used for training
- Published IP ranges
- mistral.ai/mistralai-index-ips.json
- Crawler
MistralAI-User- Operator
- Mistral
- What it is for
- Live fetch at a user’s request in Le Chat
- Published IP ranges
- mistral.ai/mistralai-user-ips.json
- Crawler
MistralAI-Training- Operator
- Mistral
- What it is for
- Training crawl only; not used for index or live answers
- Published IP ranges
- no published ranges — a request wearing this name cannot be checked
- Crawler
OAI-AdsBot- Operator
- OpenAI
- What it is for
- Validates pages submitted as ChatGPT ads; visits only submitted pages
- Published IP ranges
- openai.com/adsbot.json
- Crawler
Google-Agent- Operator
- What it is for
- User-triggered agent that navigates and acts on the web from Google infrastructure
- Published IP ranges
- developers.google.com/static/crawling/ipranges/user-triggered-agents.json
- Crawler
Google-GeminiNotebook- Operator
- What it is for
- Fetches URLs a Gemini Notebook user supplied (replaces Google-NotebookLM)
- Published IP ranges
- developers.google.com/static/crawling/ipranges/user-triggered-fetchers-google.json
- Crawler
GoogleOther-Image- Operator
- What it is for
- Google common crawler for public image URLs, governed by the GoogleOther family
- Published IP ranges
- developers.google.com/static/crawling/ipranges/common-crawlers.json
- Crawler
GoogleOther-Video- Operator
- What it is for
- Google common crawler for public video URLs, governed by the GoogleOther family
- Published IP ranges
- developers.google.com/static/crawling/ipranges/common-crawlers.json
- Crawler
ExaSearchBot- Operator
- Exa
- What it is for
- Index crawl for Exa’s AI-oriented search and retrieval; honours robots.txt
- Published IP ranges
- Web Bot Auth — every request is cryptographically signed; directory at crawler.exa.ai/.well-known/http-message-signatures-directory
Named agents per purpose, counted from this page’s own tables — 114 in total
Every agent this scanner checks, grouped by purpose
Six groups, and the grouping is the point: an agent that trains a model, one that builds a search index and one that fetches your page at the moment a person asks a question are three different things with three different consequences. The third column is the one to read — it says whether a request wearing that name can be checked at all.
Answer engines 17
These decide whether a chatbot can reach you, read you and quote you. Refusing one of these is the only kind of block that costs you answers.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
Meta-ExternalFetcher | Meta AI live fetch | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
MistralAI-User | Le Chat live fetch | answer-engine crawler | mistral.ai/mistralai-user-ips.json |
DuckAssistBot | DuckDuckGo DuckAssist | answer-engine crawler | published IP ranges |
YouBot | You.com index | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
OAI-SearchBot | ChatGPT search index | answer-engine crawler | openai.com/searchbot.json |
ChatGPT-User | ChatGPT live fetch | user-triggered fetch | openai.com/chatgpt-user.json |
Claude-SearchBot | Claude search index | answer-engine crawler | published IP ranges |
Claude-User | Claude live fetch | answer-engine crawler | published IP ranges |
PerplexityBot | Perplexity index | answer-engine crawler | published IP ranges |
Perplexity-User | Perplexity live fetch | user-triggered fetch | published IP ranges |
Gemini-Deep-Research | Gemini research agent | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
PhindBot | Phind | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
Kagibot | Kagi | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
Copilot-User | Microsoft Copilot fetch | answer-engine crawler · unverifiable | not published — a request wearing this name cannot be checked |
Amzn-User | Amazon live fetch for Alexa questions | answer-engine crawler | developer.amazon.com/amazonbot/live-ip-addresses |
Google-Agent | Google user-triggered agent, acts on the web | user-triggered fetch | user-triggered-agents.json, plus Web Bot Auth on part of its traffic |
Google-GeminiNotebook | Gemini Notebook user-supplied URL fetch | answer-engine crawler | user-triggered-fetchers-google.json (Google-owned user-triggered fetchers) |
Search indexes 24
Classic crawling, and the source most AI answer surfaces still draw on.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
Googlebot | Google Search / AI Overviews | search index crawler | reverse DNS to googlebot.com |
bingbot | Bing / Copilot | search index crawler | reverse DNS to search.msn.com |
Applebot | Apple / Siri | search index crawler | reverse DNS to applebot.apple.com |
Amazonbot | Amazon | search index crawler | developer.amazon.com/amazonbot/ip-addresses |
Googlebot-Image | Google image crawler | search index crawler | Google common-crawler ranges (shared feed) + reverse DNS |
Googlebot-News | Google News | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Googlebot-Video | Google video crawler | search index crawler | Google common-crawler ranges (shared feed) + reverse DNS |
Storebot-Google | Google Shopping | search index crawler | Google special-crawler ranges (shared feed) + reverse DNS |
Mediapartners-Google | Google AdSense | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
APIs-Google | Google push delivery | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
BingPreview | Bing page preview | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
msnbot | Microsoft legacy crawler | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Slurp | Yahoo Slurp | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Yandex | Yandex (umbrella token) | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
coccocbot-web | Coc Coc | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
DuckDuckBot | DuckDuckGo | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
YandexBot | Yandex | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Baiduspider | Baidu | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Seznambot | Seznam | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Neevabot | Neeva | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
PetalBot | Huawei Petal | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Amzn-SearchBot | Amazon search eligibility, separate from Amazonbot | search index crawler | developer.amazon.com/amazonbot/searchbot-ip-addresses |
MistralAI-Index | Mistral search index for Vibe | search index crawler | mistral.ai/mistralai-index-ips.json |
ExaSearchBot | Exa AI search index, Web Bot Auth signed | search index crawler | Web Bot Auth — signed requests, directory at crawler.exa.ai |
Training crawlers 46
They collect pages to train models. Refusing them is a policy choice with no effect on whether you appear in answers — we report it and never score it.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
cohere-ai | Cohere | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Diffbot | Diffbot knowledge graph | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Timpibot | Timpi index | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Google-CloudVertexBot | Vertex AI grounding | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
GPTBot | OpenAI training + index | training crawler | openai.com/gptbot.json |
ClaudeBot | Anthropic training | training crawler | published IP ranges |
Google-Extended | Gemini training | robots.txt control token (never sent) · unverifiable | control token only — never appears in a request |
Applebot-Extended | Apple training | robots.txt control token (never sent) · unverifiable | reverse DNS to applebot.apple.com |
CCBot | Common Crawl | training crawler | index.commoncrawl.org/ccbot.json, and reverse DNS for IPv4 |
Bytespider | ByteDance | training crawler · unverifiable | no published ranges |
meta-externalagent | Meta AI | training crawler | published IP ranges |
Magpie-crawler | Magpie AI | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
img2dataset | img2dataset image corpus | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
AwarioRssBot | Awario RSS | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
AwarioSmartBot | Awario smart | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
TurnitinBot | Turnitin | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
archive.org_bot | Internet Archive | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
ia_archiver | Internet Archive (legacy) | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
meta-webindexer | Meta web index | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
omgili | Webz.io omgili | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
cohere-training-data-crawler | Cohere training corpus | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
PanguBot | Huawei PanGu | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Ai2Bot-Dolma | Allen Institute Dolma | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
FriendlyCrawler | FriendlyCrawler ML | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
VelenPublicWebCrawler | Velen | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
MyCentralAIScraperBot | MyCentral AI | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
DeepSeekBot | DeepSeek | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
ICC-Crawler | NICT ICC | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
GoogleOther | Google non-search fetch | training crawler | Google common-crawler ranges (shared feed) + reverse DNS |
anthropic-ai | Anthropic legacy agent | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Claude-Web | Anthropic legacy agent | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
FacebookBot | Meta legacy | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Omgilibot | Webz.io | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
ImagesiftBot | Imagesift | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
AI2Bot | Allen Institute | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Scrapy | Generic scraper framework | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
SemrushBot-OCOB | Semrush AI corpus | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Applebot-Extended-Ads | Apple ads corpus | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
TikTokSpider | TikTok | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
QuillBot | QuillBot | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
Webzio-Extended | Webz.io extended | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
ProRataInc | ProRata | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
AwarioBot | Awario | training crawler · unverifiable | not published — a request wearing this name cannot be checked |
MistralAI-Training | Mistral training crawl | training crawler · unverifiable | no published ranges |
GoogleOther-Image | Google common crawler, public image URLs | training crawler | Google common-crawler ranges (shared feed) + reverse DNS |
GoogleOther-Video | Google common crawler, public video URLs | training crawler | Google common-crawler ranges (shared feed) + reverse DNS |
Regional search engines 7
Worth allowing if you serve those markets, harmless otherwise. Never scored.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
Yeti | Naver | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
YoudaoBot | Youdao | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Exabot | Exalead | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Sogou web spider | Sogou | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
YisouSpider | Yisou | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
360Spider | 360 Search | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Sosospider | Soso | search index crawler · unverifiable | not published — a request wearing this name cannot be checked |
Social link previews 6
They render the card when someone shares your URL. Never scored.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
Facebot | Meta link crawler | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked |
Twitterbot | X link preview | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked |
LinkedInBot | LinkedIn link preview | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked |
facebookexternalhit | Facebook link preview | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked |
Pinterestbot | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked | |
Slackbot-LinkExpanding | Slack unfurl | link-preview fetcher · unverifiable | not published — a request wearing this name cannot be checked |
SEO and research crawlers 14
Third-party tools. Blocking them affects nobody’s answers, including yours. Never scored.
| User-agent | What it is | Identity | How to verify it |
|---|---|---|---|
AhrefsBot | Ahrefs link index | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
SemrushBot | Semrush crawler | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
MJ12bot | Majestic link index | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
DotBot | Moz link index | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
rogerbot | Moz site crawler | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
DataForSeoBot | DataForSEO | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
BLEXBot | WebMeUp link index | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
CloudflareBrowserRenderingCrawler | Cloudflare Browser Run /crawl | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
Cloudflare-AutoRAG | Cloudflare AutoRAG | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
Peer39_Crawler | Peer39 ad context | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
AdsBot-Google | Google Ads quality | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
AmazonAdBot | Amazon Ads | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
AdIdxBot | Microsoft Ads | tool or developer fetch · unverifiable | not published — a request wearing this name cannot be checked |
OAI-AdsBot | OpenAI ChatGPT ads page validation | tool or developer fetch | openai.com/adsbot.json |
Watchlist — 10 names not yet counted
A name reaches the tables above only with a first-party operator page, an observed request, or a verifiable directory record. These have been seen in community blocklists or product notes and fail that gate today. They are not resolved from robots.txt and not sent to any site; each graduates on its own evidence, never as a batch.
| Candidate | Operator | Likely role | What would admit it |
|---|---|---|---|
kagi-fetcher | Kagi | user-triggered fetch | Confirm on Kagi’s own documentation, separately from the listed Kagibot |
iaskspider | iAsk | search index | Operator documentation or an observed full user-agent; only third-party directories name it today |
QwenBot | Alibaba / Qwen | AI crawler | First-party purpose and control documentation |
ERNIEBot | Baidu | AI crawler | First-party purpose and verification documentation |
DoubaoBot | ByteDance / Doubao | AI crawler | First-party documentation, and a stated distinction from Bytespider |
Bravebot | Brave Search | search index | Brave’s crawler page exists but does not name the token in readable form; needs the token confirmed first-party or observed in logs |
LinerBot | LINER | answer engine | getliner.com/linerbot returned 410 on 2026-09-16; needs a live first-party page |
FirecrawlAgent | Firecrawl | developer-operated fetch client | A decision on whether developer-run fetch clients belong beside operator-run crawlers |
Claude-Code | Anthropic (developer agent) | developer-agent fetch | A distinct developer-agents category; not equivalent to Anthropic’s centrally run search crawler |
Google-Gemini-CLI | Google (developer agent) | developer-agent fetch | A stable server-side identity and verification method from Google; the CLI is a local open-source agent |
What the tables cannot tell you
Blocking GPTBot does not remove you from ChatGPT
Blocking GPTBot does not remove you from ChatGPT. GPTBot collects training data. OAI-SearchBot builds the index ChatGPT search reads, and ChatGPT-User fetches your page live when someone asks about you. Those are three separate permissions, and a single Disallow aimed at the wrong one either gives away training data you meant to keep or removes you from answers you meant to appear in.
The same split applies to Google (Google-Extended is Gemini training only — it has no effect on Search or AI Overviews) and to Apple (Applebot-Extended is training only).
The user-agent string is a claim, not a fact
Anyone can send a request calling itself GPTBot. Verification means matching the request against the address ranges the operator publishes — OpenAI, Anthropic and Perplexity all publish theirs; Common Crawl and ByteDance do not.
An access log records the string. It does not record whether the string was earned. That is the difference between GPTBot visited you and something called itself GPTBot, and on our own sites the same name has arrived both ways in a single day.
Being allowed is not the same as being served
A permission in robots.txt is a request, not a delivery. Your edge, your CDN or a bot-defence rule can refuse a named crawler regardless of what your file says — and it usually does so silently, because nothing in your analytics reports a crawler that never arrived.
We have measured sites returning 502 to GPTBot and ClaudeBot while serving Googlebot normally, and a robots.txt answering HTTP 200 with a verification page that no crawler could parse. Neither shows up in a rank tracker, a validator or an uptime check.
How to check your own site
The fastest version takes ten seconds and needs nothing installed:
curl -sI -A "GPTBot/1.2" https://yoursite.com/robots.txt
Look at the status and the content type. A 200 that returns text/html means something is answering in place of your file.
For the full comparison — every agent above, what each one received, and how much of your page was readable text — run a free scan or read a real report first.
Related
- Which AI crawlers actually visit a small business site — observed visits, not probes
- The file that returned 200 to every crawler and could not be read
- The dataset — what we have measured across every site scanned here