CrawlCheck

Reference

AI crawlers: who sends them, what each one is for, and how to verify them

In one sentence

This is the reference list of the 114 named crawler agents CrawlCheck resolves from a site’s robots.txt and probes for, grouped by what each one is for — answer engines that live-fetch pages to quote them, search indexers, training-data collectors, regional search, social previews and SEO tools — with the operator, the user-agent string, and how to verify a request against the IP ranges the operator publishes.

Every crawler this scanner recognises — 114 policy identities resolved from robots.txt, of which 12 are actively sent to your homepage on every scan — grouped by what they are actually for. The same company usually sends three different crawlers for three different purposes, and blocking one does not do what people expect the others to do.

This page is the reference. The measurement is separate: a scan sends 15 client identities to your homepage from one address inside one second — GPTBot, ClaudeBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot, Amazonbot, Bytespider, Meta-ExternalAgent, CCBot, plus an unnamed client and two mobile-browser controls; the other named agents below are resolved from your robots.txt, not sent — and records what each one was actually served. Same second, same address, so a difference between two of them is your edge deciding, not the internet being busy. It also resolves every named agent in your own robots.txt into what it is permitted to do, which is how a policy someone wrote two years ago turns out to say something they did not intend.

Quick facts for every named crawler

Operator, purpose, and whether that operator publishes the IP ranges that make a claimed identity checkable. A crawler with no published feed is unverifiable here — never counted as forged.

Crawler
GPTBot
Operator
OpenAI
What it is for
Training and index crawl for ChatGPT
Published IP ranges
openai.com/gptbot.json
Crawler
OAI-SearchBot
Operator
OpenAI
What it is for
Search index behind ChatGPT answers
Published IP ranges
openai.com/searchbot.json
Crawler
ChatGPT-User
Operator
OpenAI
What it is for
Fetches a page because a user asked for it in a chat
Published IP ranges
openai.com/chatgpt-user.json
Crawler
ClaudeBot
Operator
Anthropic
What it is for
Index and training crawl
Published IP ranges
claude.com/crawling/bots.json
Crawler
Claude-User
Operator
Anthropic
What it is for
Fetches a page on a user's request
Published IP ranges
claude.com/crawling/bots.json
Crawler
Claude-SearchBot
Operator
Anthropic
What it is for
Search retrieval
Published IP ranges
claude.com/crawling/bots.json
Crawler
PerplexityBot
Operator
Perplexity
What it is for
Index crawl
Published IP ranges
www.perplexity.ai/perplexitybot.json
Crawler
Perplexity-User
Operator
Perplexity
What it is for
User-triggered fetch
Published IP ranges
www.perplexity.ai/perplexity-user.json
Crawler
Googlebot
Operator
Google
What it is for
Search index crawl
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
Crawler
Google-Extended
Operator
Google
What it is for
Training permission token, not a separate crawler
Published IP ranges
robots.txt control token only — it never appears in a request, so there is no request to verify
Crawler
Bingbot
Operator
Microsoft
What it is for
Search index crawl
Published IP ranges
www.bing.com/toolbox/bingbot.json
Crawler
Applebot
Operator
Apple
What it is for
Search and Siri index crawl
Published IP ranges
search.developer.apple.com/applebot.json
Crawler
Applebot-Extended
Operator
Apple
What it is for
Training permission token, not a separate crawler
Published IP ranges
search.developer.apple.com/applebot.json
Crawler
Amazonbot
Operator
Amazon
What it is for
Alexa and shopping index crawl
Published IP ranges
developer.amazon.com/amazonbot/ip-addresses/
Crawler
Meta-ExternalAgent
Operator
Meta
What it is for
Training and preview fetch
Published IP ranges
no JSON feed — Meta publishes its crawler ranges as ASN AS32934 (whois -h whois.radb.net -- '-i origin AS32934')
Crawler
Bytespider
Operator
ByteDance
What it is for
Training crawl
Published IP ranges
none published — claims from this identity are unverifiable
Crawler
CCBot
Operator
Common Crawl
What it is for
Open crawl corpus many models train on
Published IP ranges
index.commoncrawl.org/ccbot.json
Crawler
Amzn-SearchBot
Operator
Amazon
What it is for
Amazon search eligibility (Alexa and shopping search); not used for generative-AI training
Published IP ranges
developer.amazon.com/amazonbot/searchbot-ip-addresses/
Crawler
Amzn-User
Operator
Amazon
What it is for
Live fetch for a user action such as an Alexa question
Published IP ranges
developer.amazon.com/amazonbot/live-ip-addresses/
Crawler
MistralAI-Index
Operator
Mistral
What it is for
Index crawl for Mistral search and Vibe answers; not used for training
Published IP ranges
mistral.ai/mistralai-index-ips.json
Crawler
MistralAI-User
Operator
Mistral
What it is for
Live fetch at a user’s request in Le Chat
Published IP ranges
mistral.ai/mistralai-user-ips.json
Crawler
MistralAI-Training
Operator
Mistral
What it is for
Training crawl only; not used for index or live answers
Published IP ranges
no published ranges — a request wearing this name cannot be checked
Crawler
OAI-AdsBot
Operator
OpenAI
What it is for
Validates pages submitted as ChatGPT ads; visits only submitted pages
Published IP ranges
openai.com/adsbot.json
Crawler
Google-Agent
Operator
Google
What it is for
User-triggered agent that navigates and acts on the web from Google infrastructure
Published IP ranges
developers.google.com/static/crawling/ipranges/user-triggered-agents.json
Crawler
Google-GeminiNotebook
Operator
Google
What it is for
Fetches URLs a Gemini Notebook user supplied (replaces Google-NotebookLM)
Published IP ranges
developers.google.com/static/crawling/ipranges/user-triggered-fetchers-google.json
Crawler
GoogleOther-Image
Operator
Google
What it is for
Google common crawler for public image URLs, governed by the GoogleOther family
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
Crawler
GoogleOther-Video
Operator
Google
What it is for
Google common crawler for public video URLs, governed by the GoogleOther family
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
Crawler
ExaSearchBot
Operator
Exa
What it is for
Index crawl for Exa’s AI-oriented search and retrieval; honours robots.txt
Published IP ranges
Web Bot Auth — every request is cryptographically signed; directory at crawler.exa.ai/.well-known/http-message-signatures-directory

Named agents per purpose, counted from this page’s own tables — 114 in total

Answer17Search24Training46Regional7Social6SEO14Answer17Search24Training46Regional7Social6SEO14

Every agent this scanner checks, grouped by purpose

Six groups, and the grouping is the point: an agent that trains a model, one that builds a search index and one that fetches your page at the moment a person asks a question are three different things with three different consequences. The third column is the one to read — it says whether a request wearing that name can be checked at all.

Answer engines 17

These decide whether a chatbot can reach you, read you and quote you. Refusing one of these is the only kind of block that costs you answers.

User-agentWhat it isIdentityHow to verify it
Meta-ExternalFetcherMeta AI live fetchanswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
MistralAI-UserLe Chat live fetchanswer-engine crawlermistral.ai/mistralai-user-ips.json
DuckAssistBotDuckDuckGo DuckAssistanswer-engine crawlerpublished IP ranges
YouBotYou.com indexanswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
OAI-SearchBotChatGPT search indexanswer-engine crawleropenai.com/searchbot.json
ChatGPT-UserChatGPT live fetchuser-triggered fetchopenai.com/chatgpt-user.json
Claude-SearchBotClaude search indexanswer-engine crawlerpublished IP ranges
Claude-UserClaude live fetchanswer-engine crawlerpublished IP ranges
PerplexityBotPerplexity indexanswer-engine crawlerpublished IP ranges
Perplexity-UserPerplexity live fetchuser-triggered fetchpublished IP ranges
Gemini-Deep-ResearchGemini research agentanswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
PhindBotPhindanswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
KagibotKagianswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
Copilot-UserMicrosoft Copilot fetchanswer-engine crawler · unverifiablenot published — a request wearing this name cannot be checked
Amzn-UserAmazon live fetch for Alexa questionsanswer-engine crawlerdeveloper.amazon.com/amazonbot/live-ip-addresses
Google-AgentGoogle user-triggered agent, acts on the webuser-triggered fetchuser-triggered-agents.json, plus Web Bot Auth on part of its traffic
Google-GeminiNotebookGemini Notebook user-supplied URL fetchanswer-engine crawleruser-triggered-fetchers-google.json (Google-owned user-triggered fetchers)

Search indexes 24

Classic crawling, and the source most AI answer surfaces still draw on.

User-agentWhat it isIdentityHow to verify it
GooglebotGoogle Search / AI Overviewssearch index crawlerreverse DNS to googlebot.com
bingbotBing / Copilotsearch index crawlerreverse DNS to search.msn.com
ApplebotApple / Sirisearch index crawlerreverse DNS to applebot.apple.com
AmazonbotAmazonsearch index crawlerdeveloper.amazon.com/amazonbot/ip-addresses
Googlebot-ImageGoogle image crawlersearch index crawlerGoogle common-crawler ranges (shared feed) + reverse DNS
Googlebot-NewsGoogle Newssearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
Googlebot-VideoGoogle video crawlersearch index crawlerGoogle common-crawler ranges (shared feed) + reverse DNS
Storebot-GoogleGoogle Shoppingsearch index crawlerGoogle special-crawler ranges (shared feed) + reverse DNS
Mediapartners-GoogleGoogle AdSensesearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
APIs-GoogleGoogle push deliverysearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
BingPreviewBing page previewsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
msnbotMicrosoft legacy crawlersearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
SlurpYahoo Slurpsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
YandexYandex (umbrella token)search index crawler · unverifiablenot published — a request wearing this name cannot be checked
coccocbot-webCoc Cocsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
DuckDuckBotDuckDuckGosearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
YandexBotYandexsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
BaiduspiderBaidusearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
SeznambotSeznamsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
NeevabotNeevasearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
PetalBotHuawei Petalsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
Amzn-SearchBotAmazon search eligibility, separate from Amazonbotsearch index crawlerdeveloper.amazon.com/amazonbot/searchbot-ip-addresses
MistralAI-IndexMistral search index for Vibesearch index crawlermistral.ai/mistralai-index-ips.json
ExaSearchBotExa AI search index, Web Bot Auth signedsearch index crawlerWeb Bot Auth — signed requests, directory at crawler.exa.ai

Training crawlers 46

They collect pages to train models. Refusing them is a policy choice with no effect on whether you appear in answers — we report it and never score it.

User-agentWhat it isIdentityHow to verify it
cohere-aiCoheretraining crawler · unverifiablenot published — a request wearing this name cannot be checked
DiffbotDiffbot knowledge graphtraining crawler · unverifiablenot published — a request wearing this name cannot be checked
TimpibotTimpi indextraining crawler · unverifiablenot published — a request wearing this name cannot be checked
Google-CloudVertexBotVertex AI groundingtraining crawler · unverifiablenot published — a request wearing this name cannot be checked
GPTBotOpenAI training + indextraining crawleropenai.com/gptbot.json
ClaudeBotAnthropic trainingtraining crawlerpublished IP ranges
Google-ExtendedGemini trainingrobots.txt control token (never sent) · unverifiablecontrol token only — never appears in a request
Applebot-ExtendedApple trainingrobots.txt control token (never sent) · unverifiablereverse DNS to applebot.apple.com
CCBotCommon Crawltraining crawlerindex.commoncrawl.org/ccbot.json, and reverse DNS for IPv4
BytespiderByteDancetraining crawler · unverifiableno published ranges
meta-externalagentMeta AItraining crawlerpublished IP ranges
Magpie-crawlerMagpie AItraining crawler · unverifiablenot published — a request wearing this name cannot be checked
img2datasetimg2dataset image corpustraining crawler · unverifiablenot published — a request wearing this name cannot be checked
AwarioRssBotAwario RSStraining crawler · unverifiablenot published — a request wearing this name cannot be checked
AwarioSmartBotAwario smarttraining crawler · unverifiablenot published — a request wearing this name cannot be checked
TurnitinBotTurnitintraining crawler · unverifiablenot published — a request wearing this name cannot be checked
archive.org_botInternet Archivetraining crawler · unverifiablenot published — a request wearing this name cannot be checked
ia_archiverInternet Archive (legacy)training crawler · unverifiablenot published — a request wearing this name cannot be checked
meta-webindexerMeta web indextraining crawler · unverifiablenot published — a request wearing this name cannot be checked
omgiliWebz.io omgilitraining crawler · unverifiablenot published — a request wearing this name cannot be checked
cohere-training-data-crawlerCohere training corpustraining crawler · unverifiablenot published — a request wearing this name cannot be checked
PanguBotHuawei PanGutraining crawler · unverifiablenot published — a request wearing this name cannot be checked
Ai2Bot-DolmaAllen Institute Dolmatraining crawler · unverifiablenot published — a request wearing this name cannot be checked
FriendlyCrawlerFriendlyCrawler MLtraining crawler · unverifiablenot published — a request wearing this name cannot be checked
VelenPublicWebCrawlerVelentraining crawler · unverifiablenot published — a request wearing this name cannot be checked
MyCentralAIScraperBotMyCentral AItraining crawler · unverifiablenot published — a request wearing this name cannot be checked
DeepSeekBotDeepSeektraining crawler · unverifiablenot published — a request wearing this name cannot be checked
ICC-CrawlerNICT ICCtraining crawler · unverifiablenot published — a request wearing this name cannot be checked
GoogleOtherGoogle non-search fetchtraining crawlerGoogle common-crawler ranges (shared feed) + reverse DNS
anthropic-aiAnthropic legacy agenttraining crawler · unverifiablenot published — a request wearing this name cannot be checked
Claude-WebAnthropic legacy agenttraining crawler · unverifiablenot published — a request wearing this name cannot be checked
FacebookBotMeta legacytraining crawler · unverifiablenot published — a request wearing this name cannot be checked
OmgilibotWebz.iotraining crawler · unverifiablenot published — a request wearing this name cannot be checked
ImagesiftBotImagesifttraining crawler · unverifiablenot published — a request wearing this name cannot be checked
AI2BotAllen Institutetraining crawler · unverifiablenot published — a request wearing this name cannot be checked
ScrapyGeneric scraper frameworktraining crawler · unverifiablenot published — a request wearing this name cannot be checked
SemrushBot-OCOBSemrush AI corpustraining crawler · unverifiablenot published — a request wearing this name cannot be checked
Applebot-Extended-AdsApple ads corpustraining crawler · unverifiablenot published — a request wearing this name cannot be checked
TikTokSpiderTikToktraining crawler · unverifiablenot published — a request wearing this name cannot be checked
QuillBotQuillBottraining crawler · unverifiablenot published — a request wearing this name cannot be checked
Webzio-ExtendedWebz.io extendedtraining crawler · unverifiablenot published — a request wearing this name cannot be checked
ProRataIncProRatatraining crawler · unverifiablenot published — a request wearing this name cannot be checked
AwarioBotAwariotraining crawler · unverifiablenot published — a request wearing this name cannot be checked
MistralAI-TrainingMistral training crawltraining crawler · unverifiableno published ranges
GoogleOther-ImageGoogle common crawler, public image URLstraining crawlerGoogle common-crawler ranges (shared feed) + reverse DNS
GoogleOther-VideoGoogle common crawler, public video URLstraining crawlerGoogle common-crawler ranges (shared feed) + reverse DNS

Regional search engines 7

Worth allowing if you serve those markets, harmless otherwise. Never scored.

User-agentWhat it isIdentityHow to verify it
YetiNaversearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
YoudaoBotYoudaosearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
ExabotExaleadsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
Sogou web spiderSogousearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
YisouSpiderYisousearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
360Spider360 Searchsearch index crawler · unverifiablenot published — a request wearing this name cannot be checked
SosospiderSososearch index crawler · unverifiablenot published — a request wearing this name cannot be checked

Social link previews 6

They render the card when someone shares your URL. Never scored.

User-agentWhat it isIdentityHow to verify it
FacebotMeta link crawlerlink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked
TwitterbotX link previewlink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked
LinkedInBotLinkedIn link previewlink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked
facebookexternalhitFacebook link previewlink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked
PinterestbotPinterestlink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked
Slackbot-LinkExpandingSlack unfurllink-preview fetcher · unverifiablenot published — a request wearing this name cannot be checked

SEO and research crawlers 14

Third-party tools. Blocking them affects nobody’s answers, including yours. Never scored.

User-agentWhat it isIdentityHow to verify it
AhrefsBotAhrefs link indextool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
SemrushBotSemrush crawlertool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
MJ12botMajestic link indextool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
DotBotMoz link indextool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
rogerbotMoz site crawlertool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
DataForSeoBotDataForSEOtool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
BLEXBotWebMeUp link indextool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
CloudflareBrowserRenderingCrawlerCloudflare Browser Run /crawltool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
Cloudflare-AutoRAGCloudflare AutoRAGtool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
Peer39_CrawlerPeer39 ad contexttool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
AdsBot-GoogleGoogle Ads qualitytool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
AmazonAdBotAmazon Adstool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
AdIdxBotMicrosoft Adstool or developer fetch · unverifiablenot published — a request wearing this name cannot be checked
OAI-AdsBotOpenAI ChatGPT ads page validationtool or developer fetchopenai.com/adsbot.json

Watchlist — 10 names not yet counted

A name reaches the tables above only with a first-party operator page, an observed request, or a verifiable directory record. These have been seen in community blocklists or product notes and fail that gate today. They are not resolved from robots.txt and not sent to any site; each graduates on its own evidence, never as a batch.

CandidateOperatorLikely roleWhat would admit it
kagi-fetcherKagiuser-triggered fetchConfirm on Kagi’s own documentation, separately from the listed Kagibot
iaskspideriAsksearch indexOperator documentation or an observed full user-agent; only third-party directories name it today
QwenBotAlibaba / QwenAI crawlerFirst-party purpose and control documentation
ERNIEBotBaiduAI crawlerFirst-party purpose and verification documentation
DoubaoBotByteDance / DoubaoAI crawlerFirst-party documentation, and a stated distinction from Bytespider
BravebotBrave Searchsearch indexBrave’s crawler page exists but does not name the token in readable form; needs the token confirmed first-party or observed in logs
LinerBotLINERanswer enginegetliner.com/linerbot returned 410 on 2026-09-16; needs a live first-party page
FirecrawlAgentFirecrawldeveloper-operated fetch clientA decision on whether developer-run fetch clients belong beside operator-run crawlers
Claude-CodeAnthropic (developer agent)developer-agent fetchA distinct developer-agents category; not equivalent to Anthropic’s centrally run search crawler
Google-Gemini-CLIGoogle (developer agent)developer-agent fetchA stable server-side identity and verification method from Google; the CLI is a local open-source agent

What the tables cannot tell you

Blocking GPTBot does not remove you from ChatGPT

Blocking GPTBot does not remove you from ChatGPT. GPTBot collects training data. OAI-SearchBot builds the index ChatGPT search reads, and ChatGPT-User fetches your page live when someone asks about you. Those are three separate permissions, and a single Disallow aimed at the wrong one either gives away training data you meant to keep or removes you from answers you meant to appear in.

The same split applies to Google (Google-Extended is Gemini training only — it has no effect on Search or AI Overviews) and to Apple (Applebot-Extended is training only).

The user-agent string is a claim, not a fact

Anyone can send a request calling itself GPTBot. Verification means matching the request against the address ranges the operator publishes — OpenAI, Anthropic and Perplexity all publish theirs; Common Crawl and ByteDance do not.

An access log records the string. It does not record whether the string was earned. That is the difference between GPTBot visited you and something called itself GPTBot, and on our own sites the same name has arrived both ways in a single day.

Being allowed is not the same as being served

A permission in robots.txt is a request, not a delivery. Your edge, your CDN or a bot-defence rule can refuse a named crawler regardless of what your file says — and it usually does so silently, because nothing in your analytics reports a crawler that never arrived.

We have measured sites returning 502 to GPTBot and ClaudeBot while serving Googlebot normally, and a robots.txt answering HTTP 200 with a verification page that no crawler could parse. Neither shows up in a rank tracker, a validator or an uptime check.

How to check your own site

The fastest version takes ten seconds and needs nothing installed:

curl -sI -A "GPTBot/1.2" https://yoursite.com/robots.txt

Look at the status and the content type. A 200 that returns text/html means something is answering in place of your file.

For the full comparison — every agent above, what each one received, and how much of your page was readable text — run a free scan or read a real report first.

Related

Questions about this page

QWhat is the difference between an answer-engine crawler and a training crawler?
An answer-engine crawler (OAI-SearchBot, Claude-SearchBot, PerplexityBot) fetches pages to cite them in an answer now. A training crawler (GPTBot, ClaudeBot, CCBot) collects pages to train future models. Blocking the second while allowing the first is a coherent policy, not a defect, and the scanner treats it that way.
QWhich AI crawlers publish IP ranges I can verify against?
OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot and its agents), Perplexity, Google, Bing and Apple publish machine-readable ranges. Amazonbot and Meta-ExternalAgent do not, so a request claiming them can be neither confirmed nor accused.
QShould I block AI crawlers in robots.txt?
That is a policy choice, not a measurement. What the scanner reports is whether the rules you wrote actually apply: a named User-agent group with no Disallow lines means allow everything for that crawler, and a robots.txt served as HTML applies to nobody.
QHow often is this list updated?
As operators publish changes to their agents and ranges. The scanner’s agent table and this page are generated from the same array, so they cannot disagree.

Keep reading