CrawlCheck

Guides · 2026-10-03 · By · 0 views

Best AI web crawler on a budget: real costs compared

Five budget crawlers, their published prices converted into the pages $19 actually buys, the proxy and retry costs no pricing page shows, and the one question none of them can answer: whether ChatGPT and Claude can read your own site.

On a limited budget, pick by job. For clean Markdown into an LLM pipeline, Firecrawl Hobby ($19/month, 5,000 one-credit pages). For many JavaScript-rendered pages by API, ScrapingBee Hobby ($19, 75,000 credits at 5 per rendered page, about 15,000 pages). For scheduled scrapers and proxies, Apify Starter ($19 of usage). For zero subscription with a developer, Crawl4AI or Scrapy. To check whether ChatGPT, Claude and Perplexity can read your own site, none of these apply; run a free CrawlCheck scan.

Share of all scans carrying each finding named abovePAGE_IS_MOSTLY_CODE29.2%ANSWER_ENGINE_REFUSED5.9%AI_OPTOUT_SET5%CONTENT_NEEDS_JAVASCRIPT1.1%CHALLENGE_SERVED_2000.6%CRAWLER_SERVED_LESS0.4%Share of all scans carrying eachfinding named abovePAGE_IS_MOSTLY_CODE29.2%ANSWER_ENGINE_REFUSED5.9%AI_OPTOUT_SET5%CONTENT_NEEDS_JAVASCRIPT1.1%CHALLENGE_SERVED_2000.6%CRAWLER_SERVED_LESS0.4%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Most budget crawler comparisons stop at the sticker price, and the sticker price is the least useful number on the page. Every plan below costs about $19 a month. What that $19 buys ranges from roughly 1,000 to 15,000 JavaScript-rendered pages, depending on how each provider counts a page. This guide puts every number next to the published page it came from (all read on 3 October 2026), shows the arithmetic, and then answers the question most site owners actually came with: whether AI systems like ChatGPT and Claude can read their own site. A scraper cannot answer that one, and the reason is in section six.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

The short answer: which AI web crawler fits a limited budget #

Pick by job, not by price. All five tools below are real, maintained and honest about what they do; they do different jobs.

Your jobBest fit on a budgetWhy
Clean Markdown for an LLM or RAG pipelineFirecrawl Hobby ($19/mo, $16 on annual)5,000 credits; a page costs 1 credit, 5 with JSON extraction
Lots of JavaScript-heavy pages through an APIScrapingBee Hobby ($19/mo)75,000 credits; a JavaScript-rendered request costs 5, so about 15,000 pages
Scheduled scrapers, proxies and a marketplace of ready-made actorsApify Starter ($19/mo)$19 of platform usage included, compute at $0.20 per unit, residential proxy $8/GB
No-code monitoring of a page for changesBrowse AI Personal ($19/mo billed annually)1,000 credits a month; a credit extracts 10 rows or one screenshot
Zero subscription, full control, a developer on handCrawl4AI or Scrapy (free, open source)You pay in servers and maintenance time instead
Checking whether ChatGPT, Claude and Perplexity can read your siteNone of the above: a crawler audit (CrawlCheck, free scan)A scraper sees what the scraper is served, not what GPTBot or ClaudeBot is served

Why "credits" are not pages: the real cost per JavaScript page #

Each provider meters differently, so two $19 plans are not comparable until you convert credits into the pages you actually need. The conversion depends on whether the page needs a real browser to render (single-page apps, most modern storefronts) and whether the site sits behind bot protection that needs a premium proxy.

PlanIncludedPlain HTML pageJavaScript-rendered pageBehind bot protection
Firecrawl Hobby5,000 credits1 credit → 5,000 pages1 credit → 5,000 pages (pricing lists no separate render charge); JSON extraction +4 → 1,000 pagesNot priced separately on the pricing page
ScrapingBee Hobby75,000 credits1 credit → 75,0005 credits (the default) → 15,000Premium proxy + JS 25 → 3,000; stealth 75 → 1,000
Apify Starter$19 usage creditDepends on the actor and compute usedBrowser actors burn more compute units at $0.20 eachResidential proxy $8 per GB
Browse AI Personal (annual)1,000 credits/mo1 credit = 10 rows or one screenshotSame; premium sites cost 2 to 10 credits per taskIncluded in the premium-site rate
CrawlCheck (free scan; Watch $29/mo for 3 domains; API $99/mo)Unlimited free scans, no account; API 2,000 scans/moNot a scraper: one scan reads your whole delivery, as a browser and as each AI crawler identityMeasures the HTML delivered before JavaScript runs, which is what most AI crawlers read, and flags content that only exists after itRecords a challenge or refusal as a finding instead of routing around it

Sources: Firecrawl's pricing page, ScrapingBee's pricing page and its API documentation (which states JavaScript rendering is on by default and costs 5 credits), Apify's pricing page, Browse AI's pricing page. The quickest test before you buy anything: ask the provider how many of your pages, rendered the way you need them, the monthly allowance covers. That one number replaces every comparison table, including this one.

CrawlCheck vs Firecrawl, ScrapingBee, Apify, Browse AI and Crawl4AI #

We publish CrawlCheck, so here it is in the same table as everything else, held to the same rule: only what each product’s own page states. The honest summary is that CrawlCheck is not a cheaper scraper. It does a different job: it tells you whether AI crawlers can reach, read and quote your site. If you need to extract data from other sites, use one of the others.

ToolBuilt forOutputFreeEntry paid plan
CrawlCheckChecking whether AI crawlers (GPTBot, ClaudeBot, PerplexityBot and others) can reach, read and quote a siteGraded report with every finding ranked by severity; a signed per-domain answer by APIUnlimited scans, no account, no cardWatch $29/mo for 3 domains; API $99/mo (2,000 scans); Agency $249/mo (40 domains)
FirecrawlTurning pages into clean Markdown or JSON for LLM pipelinesMarkdown, JSON1,000 creditsHobby $19/mo ($16 annual), 5,000 credits
ScrapingBeeScraping JavaScript-heavy pages through an API with proxiesRaw HTML or extracted data1,000 API credits, no cardHobby $19/mo, 75,000 credits
ApifyRunning scheduled scrapers (actors) with proxies and storageDatasets via API$5 of usage a month, no cardStarter $19/mo of usage
Browse AINo-code extraction and change monitoringRows, screenshots50 credits a monthPersonal $19/mo annual, $48 monthly
Crawl4AISelf-hosted crawling to LLM-ready MarkdownMarkdown, structured extractionOpen source (Apache 2.0)None; you pay for servers

Where they overlap: any of the scrapers can fetch your own homepage. What none of them is built to do, and what CrawlCheck exists for, is fetch it once per AI crawler identity, compare what each one was served against a browser, read robots.txt and the AI opt-out signals per crawler, and tell you which of those differences stops ChatGPT, Claude or Perplexity from quoting you. Where CrawlCheck falls short of them: it does not extract data from other people’s sites, export datasets or run arbitrary scraping jobs, and it cannot send requests from the AI companies’ own IP addresses.

The hidden costs: proxies, bandwidth and retries #

The subscription is rarely the expensive part once a crawl gets real. Three line items grow with volume and never appear on a pricing page.

Residential proxy bandwidth. Sites with bot protection refuse datacenter addresses, so crawlers route through residential IPs billed by the gigabyte. Apify publishes $8 per GB on its Starter plan. The arithmetic is simple and worth doing before you commit: pages × megabytes per rendered page × price per gigabyte. At 2 MB a page, 10,000 pages is 20 GB, or $160 a month in proxy traffic alone, eight times the plan price. Check your own pages’ transfer size in the browser’s network panel rather than trusting any average, including ours.

Retries. A refused or timed-out request is usually retried, and on credit-metered plans a retried page can be billed twice. Watch the success rate in the provider’s dashboard for the first week; a low one means the real cost per delivered page is higher than the table above.

Rendering. A real browser is the expensive way to fetch a page, in credits on a hosted plan and in memory on your own server. Render only what needs it: many pages deliver their text in the first HTML response, and fetching those with a browser multiplies cost for no gain.

Free and open source: Crawl4AI and Scrapy #

the Crawl4AI repository is Apache-2.0 licensed, at version 0.9.4 (released 23 September 2026) with about 84,600 GitHub stars when we read it. It turns pages into LLM-ready Markdown, supports CSS, XPath, regex and LLM-based extraction, and runs real browsers through Playwright, so its setup installs Chromium. Scrapy is the long-standing Python framework for high-volume HTML crawling; it does not render JavaScript on its own and is usually paired with a browser service when pages need one.

Free software is not a free crawl. You pay for the server, for browser memory when pages need rendering, for proxies when sites refuse datacenter traffic, and for the hours spent when a site changes its markup and a selector stops matching. For a team with a developer, that trade is often worth it. For a small business without one, a hosted plan usually costs less in total.

The gap no budget crawler covers: what AI crawlers actually receive #

A general-purpose crawler tells you what that crawler was served. It does not tell you what GPTBot, ClaudeBot or PerplexityBot were served, and those can differ, because sites and their CDNs treat crawlers by name. Firecrawl, for example, identifies itself as FirecrawlAgent and applies the robots.txt rules written for that token (per Firecrawl's own documentation). A site that blocks GPTBot can serve FirecrawlAgent perfectly, and the reverse.

The AI companies run several crawlers each, for different jobs. OpenAI's crawler documentation lists GPTBot (training), OAI-SearchBot (ChatGPT search results) and ChatGPT-User (pages a user asks ChatGPT to open, where robots.txt may not apply), each with a published IP list. Anthropic's crawler page lists ClaudeBot, Claude-User and Claude-SearchBot the same way. Blocking one does not block the others, which is why a site can disappear from ChatGPT search while still being used for training, or the other way round.

CrawlerCompanyWhat it is forControlled by robots.txt
GPTBotOpenAITraining dataYes
OAI-SearchBotOpenAIResults in ChatGPT searchYes
ChatGPT-UserOpenAIA page a user asks ChatGPT to openNot necessarily
ClaudeBotAnthropicTraining dataYes
Claude-SearchBotAnthropicSearch qualityYes
Claude-UserAnthropicA page a user asks Claude to openYes, per Anthropic

What we measure across real sites #

CrawlCheck fetches each site the way each crawler identifies itself and compares what each was served. Across every scan on record, these are the shares of sites carrying each problem. The figures are read live from the dataset each time this page renders, so they move as scans arrive.

What goes wrongShare of scanned sites
An answer engine’s crawler refused while a browser was served5.9%
A named crawler served materially less text than an ordinary client0.4%
A verification (challenge) page served with status 200 instead of the content0.6%
Content that only exists after JavaScript runs1.1%
Almost nothing a machine receives is readable text29.2%
An AI opt-out signal already set5%

Live, as you read this: the corpus now holds 189,169 domains across 3,137 scans. The figures in this piece were measured on the date above; this line is not.

None of these show up in a scraper’s output, because the scraper is not the one being refused, challenged or served less. The full method and every code are on the dataset page; the JavaScript question is covered in depth in do AI crawlers render JavaScript.

How to check whether ChatGPT and Claude can read your site, free #

Run a free scan at crawlcheck.io (no account, no card). It fetches your site as a browser and as each AI crawler identifies itself, compares what each received, reads your robots.txt per crawler, and returns a grade with every finding named and ranked by severity. It sends each crawler’s published user-agent; it cannot send from the companies’ own IP addresses, so a block keyed only to their IP ranges is reported as a difference between identities rather than guessed at. Step-by-step: how to check if AI crawlers can read your site.

If you would rather not do the fixes yourself, the fix list is $749 once, implemented and re-scanned so the result is measured rather than asserted. Ongoing monitoring is $29 a month for up to three domains, and agencies get 40 domains with API access for $249 a month. Developers and agent builders can read the same answer by API from $99 a month for 2,000 scans, or free one domain at a time through the open resolve protocol.

A decision checklist before you pay for any crawler #

How these numbers were checked #

Every price and allowance above was read from the provider’s own pricing or documentation page on 3 October 2026 and is linked where it appears. Prices change; the links are the source of truth. Where a provider does not publish a figure (for example, separate JavaScript or proxy charges on a pricing page), the table says so instead of estimating. The share-of-sites figures come from CrawlCheck’s own scans and render live. Crawler purposes come from OpenAI’s and Anthropic’s own documentation. Related reading: the AI crawler list and how to check if a GPTBot request is real.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is the best AI web crawler for a limited budget?

It depends on the job. Firecrawl Hobby ($19 a month, 5,000 credits at one credit a page) suits LLM and RAG pipelines that need clean Markdown. ScrapingBee Hobby ($19, 75,000 credits) gives about 15,000 JavaScript-rendered pages at its default five credits each. Apify Starter ($19 of usage) suits scheduled scrapers with proxies. Crawl4AI and Scrapy are free and open source if you have a developer to run them.

How does CrawlCheck compare to Firecrawl or ScrapingBee?

They do different jobs. Firecrawl and ScrapingBee fetch and extract pages from any site, priced by credits from $19 a month. CrawlCheck is not a scraper: it checks whether AI crawlers such as GPTBot, ClaudeBot and PerplexityBot can reach, read and quote your own site, comparing what each crawler identity is served against a browser. Its scans are free and unlimited; monitoring is $29 a month for up to three domains and the API starts at $99 a month.

How many JavaScript pages does a $19 crawler plan really cover?

Convert credits into rendered pages. ScrapingBee charges 5 credits for a JavaScript-rendered request, so 75,000 credits is about 15,000 pages, 3,000 with a premium proxy and 1,000 with a stealth proxy. Firecrawl charges 1 credit a page (5 with JSON extraction), so 5,000 credits is 5,000 pages, or 1,000 with extraction. Prices were read from each provider's own pages on 3 October 2026.

Is Crawl4AI free?

Yes. Crawl4AI is open source under the Apache 2.0 licence and outputs LLM-ready Markdown. It runs real browsers through Playwright, so you pay for the server, its memory, any proxies and the time to maintain it rather than a subscription.

Does a web scraper see what GPTBot or ClaudeBot see?

No. Each crawler is served according to its own name, and sites and CDNs often treat crawlers differently. Firecrawl identifies itself as FirecrawlAgent, for example, so a site that blocks GPTBot can still serve it, and the reverse. Testing your scraper's access says nothing about whether ChatGPT or Claude can read the site.

How do I check if ChatGPT and Claude can read my website?

Run a free scan at crawlcheck.io. It fetches the site as a browser and as each AI crawler identifies itself, compares what each was served, reads robots.txt per crawler and lists every problem ranked by severity, with no account or card. It sends each crawler's published user-agent but cannot send from the AI companies' own IP addresses.

What hidden costs make a cheap crawler expensive?

Residential proxy bandwidth (Apify publishes $8 per GB, so 10,000 two-megabyte pages is about $160 a month), retried requests that are billed again, browser rendering that multiplies credit or memory use, and the maintenance time when a site changes and a selector stops matching.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works