CrawlCheck

Reference

AI crawlers: who sends them, what each one is for, and how to verify them

In one sentence

This is the reference list of the 92 named crawler agents CrawlCheck resolves from a site’s robots.txt and probes for, grouped by what each one is for — answer engines that live-fetch pages to quote them, search indexers, training-data collectors, regional search, social previews and SEO tools — with the operator, the user-agent string, and how to verify a request against the IP ranges the operator publishes.

Every crawler this scanner checks for — 92 named agents, grouped by what they are actually for. The same company usually sends three different crawlers for three different purposes, and blocking one does not do what people expect the others to do.

This page is the reference. The measurement is separate: a scan sends fifteen client identities to your homepage from one address inside one second — the named crawlers below plus a plain browser and a mobile browser as controls — and records what each one was actually served. Same second, same address, so a difference between two of them is your edge deciding, not the internet being busy. It also resolves every named agent in your own robots.txt into what it is permitted to do, which is how a policy someone wrote two years ago turns out to say something they did not intend.

Quick facts for every named crawler

Operator, purpose, and whether that operator publishes the IP ranges that make a claimed identity checkable. A crawler with no published feed is unverifiable here — never counted as forged.

Crawler
GPTBot
Operator
OpenAI
What it is for
Training and index crawl for ChatGPT
Published IP ranges
openai.com/gptbot.json
Crawler
OAI-SearchBot
Operator
OpenAI
What it is for
Search index behind ChatGPT answers
Published IP ranges
openai.com/searchbot.json
Crawler
ChatGPT-User
Operator
OpenAI
What it is for
Fetches a page because a user asked for it in a chat
Published IP ranges
openai.com/chatgpt-user.json
Crawler
ClaudeBot
Operator
Anthropic
What it is for
Index and training crawl
Published IP ranges
claude.com/crawling/bots.json
Crawler
Claude-User
Operator
Anthropic
What it is for
Fetches a page on a user's request
Published IP ranges
claude.com/crawling/bots.json
Crawler
Claude-SearchBot
Operator
Anthropic
What it is for
Search retrieval
Published IP ranges
claude.com/crawling/bots.json
Crawler
PerplexityBot
Operator
Perplexity
What it is for
Index crawl
Published IP ranges
www.perplexity.ai/perplexitybot.json
Crawler
Perplexity-User
Operator
Perplexity
What it is for
User-triggered fetch
Published IP ranges
www.perplexity.ai/perplexity-user.json
Crawler
Googlebot
Operator
Google
What it is for
Search index crawl
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
Crawler
Google-Extended
Operator
Google
What it is for
Training permission token, not a separate crawler
Published IP ranges
developers.google.com/static/crawling/ipranges/common-crawlers.json
Crawler
Bingbot
Operator
Microsoft
What it is for
Search index crawl
Published IP ranges
www.bing.com/toolbox/bingbot.json
Crawler
Applebot
Operator
Apple
What it is for
Search and Siri index crawl
Published IP ranges
search.developer.apple.com/applebot.json
Crawler
Applebot-Extended
Operator
Apple
What it is for
Training permission token, not a separate crawler
Published IP ranges
search.developer.apple.com/applebot.json
Crawler
Amazonbot
Operator
Amazon
What it is for
Alexa and shopping index crawl
Published IP ranges
none published — claims from this identity are unverifiable
Crawler
Meta-ExternalAgent
Operator
Meta
What it is for
Training and preview fetch
Published IP ranges
none published — claims from this identity are unverifiable
Crawler
Bytespider
Operator
ByteDance
What it is for
Training crawl
Published IP ranges
none published — claims from this identity are unverifiable
Crawler
CCBot
Operator
Common Crawl
What it is for
Open crawl corpus many models train on
Published IP ranges
none published — claims from this identity are unverifiable

Named agents per purpose, counted from this page’s own tables — 92 in total

Answer14Search10Training43Regional7Social5SEO13

Every agent this scanner checks, grouped by purpose

Six groups, and the grouping is the point: an agent that trains a model, one that builds a search index and one that fetches your page at the moment a person asks a question are three different things with three different consequences. The third column is the one to read — it says whether a request wearing that name can be checked at all.

Answer engines 14

These decide whether a chatbot can reach you, read you and quote you. Refusing one of these is the only kind of block that costs you answers.

User-agentWhat it isHow to verify it
Meta-ExternalFetcherMeta AI live fetchnot published — a request wearing this name cannot be checked
MistralAI-UserLe Chat live fetchnot published — a request wearing this name cannot be checked
DuckAssistBotDuckDuckGo DuckAssistpublished IP ranges
YouBotYou.com indexnot published — a request wearing this name cannot be checked
OAI-SearchBotChatGPT search indexopenai.com/searchbot.json
ChatGPT-UserChatGPT live fetchopenai.com/chatgpt-user.json
Claude-SearchBotClaude search indexpublished IP ranges
Claude-UserClaude live fetchpublished IP ranges
PerplexityBotPerplexity indexpublished IP ranges
Perplexity-UserPerplexity live fetchpublished IP ranges
Gemini-Deep-ResearchGemini research agentnot published — a request wearing this name cannot be checked
PhindBotPhindnot published — a request wearing this name cannot be checked
KagibotKaginot published — a request wearing this name cannot be checked
Copilot-UserMicrosoft Copilot fetchnot published — a request wearing this name cannot be checked

Search indexes 10

Classic crawling, and the source most AI answer surfaces still draw on.

User-agentWhat it isHow to verify it
GooglebotGoogle Search / AI Overviewsreverse DNS to googlebot.com
bingbotBing / Copilotreverse DNS to search.msn.com
ApplebotApple / Sirireverse DNS to applebot.apple.com
AmazonbotAmazonpublished IP ranges
DuckDuckBotDuckDuckGonot published — a request wearing this name cannot be checked
YandexBotYandexnot published — a request wearing this name cannot be checked
BaiduspiderBaidunot published — a request wearing this name cannot be checked
SeznambotSeznamnot published — a request wearing this name cannot be checked
NeevabotNeevanot published — a request wearing this name cannot be checked
PetalBotHuawei Petalnot published — a request wearing this name cannot be checked

Training crawlers 43

They collect pages to train models. Refusing them is a policy choice with no effect on whether you appear in answers — we report it and never score it.

User-agentWhat it isHow to verify it
cohere-aiCoherenot published — a request wearing this name cannot be checked
DiffbotDiffbot knowledge graphnot published — a request wearing this name cannot be checked
TimpibotTimpi indexnot published — a request wearing this name cannot be checked
Google-CloudVertexBotVertex AI groundingnot published — a request wearing this name cannot be checked
GPTBotOpenAI training + indexopenai.com/gptbot.json
ClaudeBotAnthropic trainingpublished IP ranges
Google-ExtendedGemini trainingreverse DNS to googlebot.com
Applebot-ExtendedApple trainingreverse DNS to applebot.apple.com
CCBotCommon Crawlno published ranges
BytespiderByteDanceno published ranges
meta-externalagentMeta AIpublished IP ranges
Magpie-crawlerMagpie AInot published — a request wearing this name cannot be checked
img2datasetimg2dataset image corpusnot published — a request wearing this name cannot be checked
AwarioRssBotAwario RSSnot published — a request wearing this name cannot be checked
AwarioSmartBotAwario smartnot published — a request wearing this name cannot be checked
TurnitinBotTurnitinnot published — a request wearing this name cannot be checked
archive.org_botInternet Archivenot published — a request wearing this name cannot be checked
ia_archiverInternet Archive (legacy)not published — a request wearing this name cannot be checked
meta-webindexerMeta web indexnot published — a request wearing this name cannot be checked
omgiliWebz.io omgilinot published — a request wearing this name cannot be checked
cohere-training-data-crawlerCohere training corpusnot published — a request wearing this name cannot be checked
PanguBotHuawei PanGunot published — a request wearing this name cannot be checked
Ai2Bot-DolmaAllen Institute Dolmanot published — a request wearing this name cannot be checked
FriendlyCrawlerFriendlyCrawler MLnot published — a request wearing this name cannot be checked
VelenPublicWebCrawlerVelennot published — a request wearing this name cannot be checked
MyCentralAIScraperBotMyCentral AInot published — a request wearing this name cannot be checked
DeepSeekBotDeepSeeknot published — a request wearing this name cannot be checked
ICC-CrawlerNICT ICCnot published — a request wearing this name cannot be checked
GoogleOtherGoogle non-search fetchnot published — a request wearing this name cannot be checked
anthropic-aiAnthropic legacy agentnot published — a request wearing this name cannot be checked
Claude-WebAnthropic legacy agentnot published — a request wearing this name cannot be checked
FacebookBotMeta legacynot published — a request wearing this name cannot be checked
OmgilibotWebz.ionot published — a request wearing this name cannot be checked
ImagesiftBotImagesiftnot published — a request wearing this name cannot be checked
AI2BotAllen Institutenot published — a request wearing this name cannot be checked
ScrapyGeneric scraper frameworknot published — a request wearing this name cannot be checked
SemrushBot-OCOBSemrush AI corpusnot published — a request wearing this name cannot be checked
Applebot-Extended-AdsApple ads corpusnot published — a request wearing this name cannot be checked
TikTokSpiderTikToknot published — a request wearing this name cannot be checked
QuillBotQuillBotnot published — a request wearing this name cannot be checked
Webzio-ExtendedWebz.io extendednot published — a request wearing this name cannot be checked
ProRataIncProRatanot published — a request wearing this name cannot be checked
AwarioBotAwarionot published — a request wearing this name cannot be checked

Regional search engines 7

Worth allowing if you serve those markets, harmless otherwise. Never scored.

User-agentWhat it isHow to verify it
YetiNavernot published — a request wearing this name cannot be checked
YoudaoBotYoudaonot published — a request wearing this name cannot be checked
ExabotExaleadnot published — a request wearing this name cannot be checked
Sogou web spiderSogounot published — a request wearing this name cannot be checked
YisouSpiderYisounot published — a request wearing this name cannot be checked
360Spider360 Searchnot published — a request wearing this name cannot be checked
SosospiderSosonot published — a request wearing this name cannot be checked

Social link previews 5

They render the card when someone shares your URL. Never scored.

User-agentWhat it isHow to verify it
TwitterbotX link previewnot published — a request wearing this name cannot be checked
LinkedInBotLinkedIn link previewnot published — a request wearing this name cannot be checked
facebookexternalhitFacebook link previewnot published — a request wearing this name cannot be checked
PinterestbotPinterestnot published — a request wearing this name cannot be checked
Slackbot-LinkExpandingSlack unfurlnot published — a request wearing this name cannot be checked

SEO and research crawlers 13

Third-party tools. Blocking them affects nobody’s answers, including yours. Never scored.

User-agentWhat it isHow to verify it
AhrefsBotAhrefs link indexnot published — a request wearing this name cannot be checked
SemrushBotSemrush crawlernot published — a request wearing this name cannot be checked
MJ12botMajestic link indexnot published — a request wearing this name cannot be checked
DotBotMoz link indexnot published — a request wearing this name cannot be checked
rogerbotMoz site crawlernot published — a request wearing this name cannot be checked
DataForSeoBotDataForSEOnot published — a request wearing this name cannot be checked
BLEXBotWebMeUp link indexnot published — a request wearing this name cannot be checked
CloudflareBrowserRenderingCrawlerCloudflare Browser Run /crawlnot published — a request wearing this name cannot be checked
Cloudflare-AutoRAGCloudflare AutoRAGnot published — a request wearing this name cannot be checked
Peer39_CrawlerPeer39 ad contextnot published — a request wearing this name cannot be checked
AdsBot-GoogleGoogle Ads qualitynot published — a request wearing this name cannot be checked
AmazonAdBotAmazon Adsnot published — a request wearing this name cannot be checked
AdIdxBotMicrosoft Adsnot published — a request wearing this name cannot be checked

What the tables cannot tell you

Blocking GPTBot does not remove you from ChatGPT

Blocking GPTBot does not remove you from ChatGPT. GPTBot collects training data. OAI-SearchBot builds the index ChatGPT search reads, and ChatGPT-User fetches your page live when someone asks about you. Those are three separate permissions, and a single Disallow aimed at the wrong one either gives away training data you meant to keep or removes you from answers you meant to appear in.

The same split applies to Google (Google-Extended is Gemini training only — it has no effect on Search or AI Overviews) and to Apple (Applebot-Extended is training only).

The user-agent string is a claim, not a fact

Anyone can send a request calling itself GPTBot. Verification means matching the request against the address ranges the operator publishes — OpenAI, Anthropic and Perplexity all publish theirs; Common Crawl and ByteDance do not.

An access log records the string. It does not record whether the string was earned. That is the difference between GPTBot visited you and something called itself GPTBot, and on our own sites the same name has arrived both ways in a single day.

Being allowed is not the same as being served

A permission in robots.txt is a request, not a delivery. Your edge, your CDN or a bot-defence rule can refuse a named crawler regardless of what your file says — and it usually does so silently, because nothing in your analytics reports a crawler that never arrived.

We have measured sites returning 502 to GPTBot and ClaudeBot while serving Googlebot normally, and a robots.txt answering HTTP 200 with a verification page that no crawler could parse. Neither shows up in a rank tracker, a validator or an uptime check.

How to check your own site

The fastest version takes ten seconds and needs nothing installed:

curl -sI -A "GPTBot/1.2" https://yoursite.com/robots.txt

Look at the status and the content type. A 200 that returns text/html means something is answering in place of your file.

For the full comparison — every agent above, what each one received, and how much of your page was readable text — run a free scan or read a real report first.

Related

Questions about this page

QWhat is the difference between an answer-engine crawler and a training crawler?
An answer-engine crawler (OAI-SearchBot, Claude-SearchBot, PerplexityBot) fetches pages to cite them in an answer now. A training crawler (GPTBot, ClaudeBot, CCBot) collects pages to train future models. Blocking the second while allowing the first is a coherent policy, not a defect, and the scanner treats it that way.
QWhich AI crawlers publish IP ranges I can verify against?
OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot and its agents), Perplexity, Google, Bing and Apple publish machine-readable ranges. Amazonbot and Meta-ExternalAgent do not, so a request claiming them can be neither confirmed nor accused.
QShould I block AI crawlers in robots.txt?
That is a policy choice, not a measurement. What the scanner reports is whether the rules you wrote actually apply: a named User-agent group with no Disallow lines means allow everything for that crawler, and a robots.txt served as HTML applies to nobody.
QHow often is this list updated?
As operators publish changes to their agents and ranges. The scanner’s agent table and this page are generated from the same array, so they cannot disagree.

Keep reading