CrawlCheck

Guides · 2026-10-03 · By · 3 views

Cloudflare AI Crawl Control vs CrawlCheck: block vs measure

One controls AI crawler access at Cloudflare's edge; the other measures, from outside, what each crawler is actually served. What each does, where each is weaker, and the routine that uses both.

Cloudflare AI Crawl Control is an edge control: on sites served through Cloudflare it shows AI crawler traffic and lets you allow or block each crawler, on all plans. CrawlCheck is a measurement: it fetches any domain as each AI crawler identifies itself and reports what each was actually served, including conflicts between robots.txt and firewall rules. Use Cloudflare to set the policy and CrawlCheck, free, to confirm the result.

Cloudflare AI Crawl Control and CrawlCheck come up in the same conversations, and they are not substitutes. AI Crawl Control is a control: it runs at Cloudflare’s edge, in front of a site, and lets the owner see AI crawler traffic and allow, block or charge each crawler. CrawlCheck is a measurement: it fetches any site from outside, as each AI crawler identifies itself, and reports what each one was actually served. The control decides; the measurement confirms. Most sites on Cloudflare need both, because the commonest way a site loses AI visibility is a setting that does one thing while the owner believes it does another.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

Key points
  • Cloudflare AI Crawl Control is available on all Cloudflare plans and has four parts: manage AI crawlers, analyze AI traffic, track robots.txt, and pay per crawl (private beta), per Cloudflare’s documentation.
  • It acts only on sites served through Cloudflare. CrawlCheck measures any public domain.
  • On a Cloudflare site, AI Crawl Control, bot management and WAF rules can each decide a crawler’s request; the first one that refuses wins.
  • Across CrawlCheck’s scans, an answer engine’s crawler is refused while a browser is served on 5.9% of sites.
  • Use AI Crawl Control to set the policy and a CrawlCheck scan to confirm what each crawler is served.

What is Cloudflare AI Crawl Control? #

AI Crawl Control is Cloudflare’s set of tools for AI crawlers on a Cloudflare zone. Its documentation lists four parts. Manage AI crawlers lets the owner choose, crawler by crawler, how each may interact with the domain. Analyze AI traffic shows how AI crawlers are reaching the site’s pages. Track robots.txt checks the health of the robots.txt file and identifies crawlers that ignore its directives. Pay per crawl, in private beta, lets AI crawlers pay for access. Cloudflare describes the feature as available on all plans.

Because it runs on Cloudflare’s network, everything it does applies to real requests as they arrive. That is its strength: it sees the actual crawlers, from their actual addresses, and it can act on them before they reach the server.

What is CrawlCheck? #

CrawlCheck is an external measurement of how AI systems receive a site. A scan fetches the homepage and the machine files around it (robots.txt, sitemaps, llms.txt) as each crawler identifies itself, among them GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot and Googlebot, alongside a browser, and compares what each was served. It resolves robots.txt for 114 user-agents and checks the result against what each identity actually received. It works on any public domain, whatever the host or CDN, and it cannot block anything; it reports, and names the change that fixes each finding.

Side by side #

Cloudflare AI Crawl ControlCrawlCheck
TypeEdge control on your own Cloudflare zoneExternal measurement of any domain
SeesAI crawler requests that reach your site through CloudflareWhat each crawler identity is served when it fetches the site
Allow or block crawlersYes, per crawlerNo; it reports, and gives the exact fix
robots.txtTracks file health and which crawlers violate its directivesResolves robots.txt per crawler (114 user-agents) and compares it with what the edge served
Pay per crawlYes (private beta, per Cloudflare)No
Sites not on CloudflareNot coveredYes, any public domain
Content readable by AI (JavaScript-only text, pages that are mostly code, challenge pages)Not part of the featureYes
Real crawler traffic from real addressesYesNo; fetches from its own network under each crawler’s user-agent
PriceAvailable on all Cloudflare plansFree scans; Watch $29/mo for 3 domains

Sources: Cloudflare AI Crawl Control documentation; CrawlCheck pricing and robots.txt checker.

Why does a site on Cloudflare still refuse AI crawlers it allows? #

Because AI Crawl Control is not the only setting that decides a crawler’s request. On a Cloudflare zone, bot management (including Bot Fight Mode and bot scores) and WAF custom rules also act on requests, and any of them can refuse or challenge a crawler that AI Crawl Control and robots.txt both allow. The request stops at the first layer that says no, so robots.txt on the server is never even read.

Where an AI crawler request is decided on a Cloudflare siteTHE CRAWLERRequestGPTBot, ClaudeBot,PerplexityBot …asks for a pageCLOUDFLARE EDGE — ANY LAYER CAN SAY NOAI Crawl Controlallow, block,charge percrawlerBot managementBot Fight Mode,bot scores,challengesWAF rulescustom ruleson user-agent,IP, pathYOUR SERVEROriginrobots.txt,page, pluginsThe first layer that refuses wins. Every later layer, including robots.txt, never sees the request.AI Crawl Control can allow a crawler that a bot rule or a WAF rule still refuses or challenges.Cloudflare dashboard seeswhat the edge decided for real requestsAn outside fetch seeswhat each identity was actually served
On a Cloudflare zone, three separate settings can decide an AI crawler’s request before it reaches the site. The dashboard reports each layer’s decision; only a fetch from outside shows the combined result a crawler receives.

None of these settings is wrong on its own. Bot management exists to stop abusive automated traffic, and a WAF rule written to stop a scraper can match a legitimate crawler’s user-agent or address range by accident. The problem is that the decisions live in different screens, and no single screen shows the combined outcome for, say, ClaudeBot on your homepage. Cloudflare reports what each layer decided for real traffic; a fetch from outside shows what the crawler ended up with.

How often does it actually go wrong? #

Measured across every site CrawlCheck has scanned, these are the access failures that matter most for AI visibility. The figures update as the dataset grows:

How often AI crawler access fails on scanned sites — live share of every CrawlCheck scan, read when this page loaded
Answer-engine crawler refused, browser served5.9%
robots.txt groups silently drop rules13.4%
An AI opt-out already set5%
Challenge page served as 2000.6%
Crawler served far less text than a browser0.4%

Each bar is the share of all scanned sites with that finding. Several of these come from edge settings that nobody on the site’s own team changed.

The first bar is the one this comparison is about: a site that serves a browser normally and refuses an answer engine’s crawler. The owner loads the site, sees it working, and has no reason to look further. The fourth is subtler: a challenge page answered with status 200, which a crawler may store as if it were the page.

Two views of the same site #

Inside view versus outside viewINSIDE: CLOUDFLARE AI CRAWL CONTROLSees and actsReal requests that reached your zoneWhich crawlers, how often, which pathsWhich crawlers ignore your robots.txtAllow, block, or charge per crawlerOnly sites served through CloudflareOnly after the fact, from real trafficOUTSIDE: CRAWLCHECKMeasures and comparesA fetch as each named crawlerStatus, size and text per identityrobots.txt resolved for 114 user-agentsChallenge pages, JavaScript-only textAny public domain, any hostCannot act; reports and names the fix
The two tools look at the same site from opposite sides. Neither view contains the other, which is why a policy set inside should be confirmed from outside.

The inside view is authoritative about real traffic: Cloudflare sees the genuine GPTBot, from OpenAI’s addresses, and records what happened to it. The outside view is authoritative about outcome per identity: it asks for the page the way each crawler asks, compares the answers side by side, and checks the content each was given. The inside view cannot tell you what a crawler that has not visited yet would get; the outside view cannot tell you how many times the real crawler came.

Which should you use? #

Which tool answers which questionWhat do you need to do?Block, allow or chargefor AI crawlers on aCloudflare siteProve thatChatGPT, Claude orPerplexity can read itKeep checkingafter changing a setting,or across many hostsAI Crawl Controlin the Cloudflare dashboardCrawlCheck scanfree, any domainBoth, in that orderchange in Cloudflare,then re-measure
Pick by the job, not the brand. Control lives at the edge; proof comes from a fetch made the way the crawler makes it.

A routine that uses both #

The routine: set, measure, fix, re-measure1 Decideper crawler: train,search, user2 SetAI Crawl Controland robots.txt3 Measurescan as eachidentity4 Fixthe rule thatoverrides the policyre-measure after every Cloudflare change
A policy is only finished when a measurement confirms it. Watch repeats step 3 twice a day and alerts when a crawler’s result changes.
  1. Decide a policy per crawler. Training crawlers (GPTBot, ClaudeBot), search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-initiated fetchers (ChatGPT-User, Claude-User, Perplexity-User) are separate decisions. Blocking a training crawler does not remove a site from that company’s search results.
  2. Set it in AI Crawl Control and in robots.txt, so that the edge and the file say the same thing.
  3. Run a free CrawlCheck scan and read the result per identity: allowed by robots.txt, served by the edge, and given the same text as a browser.
  4. Fix the difference. It is usually a bot-management or WAF rule overriding the AI Crawl Control setting.
  5. Re-measure after every Cloudflare change. Watch does this twice a day and alerts when a crawler’s result changes.

What does a mismatch look like in a report? #

A typical case reads like this: robots.txt allows ClaudeBot, AI Crawl Control is set to allow it, and the scan shows the browser identity receiving the homepage with status 200 and a full page of text while ClaudeBot receives a 403, or a 200 whose body is a short challenge page. The report puts the two responses next to each other, names the identity, and points to the layer most likely responsible. From there the fix is a Cloudflare setting, and the next scan confirms it.

Where CrawlCheck is weaker #

CrawlCheck sees what it is served when it fetches the site; it does not see the real AI crawlers’ traffic, and it cannot send requests from OpenAI’s or Anthropic’s own addresses. A rule keyed only to those companies’ IP ranges therefore shows up as a difference between identities rather than a confirmed block, and a rule that allows the published ranges while refusing everyone else using the crawler’s name will refuse CrawlCheck’s fetch too. Cloudflare’s dashboard, by contrast, sees every real request that reaches the zone and can act on it. For first-party traffic evidence and enforcement, use Cloudflare; for the per-crawler outcome from outside, on any host, use CrawlCheck.

More detail: is Cloudflare blocking AI crawlers on your site?, how to block AI crawlers properly and how to verify GPTBot, ClaudeBot and PerplexityBot.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is the difference between Cloudflare AI Crawl Control and CrawlCheck?

AI Crawl Control is a Cloudflare feature that shows AI crawler traffic on your zone and lets you allow, block or charge each crawler. CrawlCheck is an external measurement that fetches any site as each AI crawler identifies itself and reports what each was served. One sets the policy; the other confirms its effect.

Is Cloudflare AI Crawl Control free?

Cloudflare's documentation says AI Crawl Control is available on all plans. Pay per crawl, which charges AI crawlers for access, is in private beta according to the same documentation.

What are the parts of Cloudflare AI Crawl Control?

Per Cloudflare's documentation: manage AI crawlers, analyze AI traffic, track robots.txt (file health and which crawlers violate its directives), and pay per crawl in private beta.

Can Cloudflare block AI crawlers that I have allowed?

Yes. On a Cloudflare zone, bot management, Bot Fight Mode and WAF custom rules also act on requests and can refuse or challenge a crawler that robots.txt and AI Crawl Control allow. An external scan that fetches as each crawler shows the combined result.

Do I need CrawlCheck if I use Cloudflare?

If you want proof of what each AI crawler is served, yes. Cloudflare reports what each layer of your edge did with real traffic; CrawlCheck checks the outcome per crawler from outside and flags conflicts between robots.txt and edge behaviour. Scans are free.

Does CrawlCheck work for sites not on Cloudflare?

Yes. It measures any public domain regardless of host or CDN. Cloudflare AI Crawl Control only applies to sites served through Cloudflare.

Can CrawlCheck block AI crawlers?

No. CrawlCheck measures and reports, including the exact setting to change. Blocking happens in robots.txt, the CDN or the firewall, such as Cloudflare AI Crawl Control.

Does blocking GPTBot in Cloudflare remove my site from ChatGPT search?

Not by itself. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the crawler for ChatGPT search, so each needs its own decision in Cloudflare and in robots.txt.

Does Cloudflare's managed robots.txt change my file?

Yes. With managed robots.txt on, Cloudflare puts its own rules before your file and serves both as one response, or creates the file if you have none. Its rules disallow Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. The file in your repository can say Allow while the live file says Disallow, so always test the live /robots.txt.

Does blocking AI training bots in Cloudflare also block Google Search?

Not through managed robots.txt. Cloudflare's documentation says it disallows Google-Extended, which governs use of your content for Google's AI models, and leaves Googlebot, which indexes for Search, allowed. A WAF custom rule you wrote yourself can still block Googlebot, so check those separately.

How does Cloudflare group AI crawlers?

AI Crawl Control lets you filter crawlers by category, such as AI crawler, AI assistant and archiver, and set Allow or Block for each crawler in the Actions column. On the free plan it identifies crawlers by user-agent string; paid plans add Cloudflare's bot detection.

How do I check whether Cloudflare is blocking a specific bot?

In AI Crawl Control, the Requests column shows allowed and unsuccessful requests for each crawler, and Security Events filtered by the bot's user agent shows which rule acted. From outside, fetch the page with that bot's user agent and compare the status and body with a normal browser fetch, which is what a CrawlCheck scan does for each named crawler. An outside fetch carries the bot's name but not its IP address, so a rule that verifies addresses can treat it differently.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works