CrawlCheck

Guides · 2026-10-03 · By · 0 views

Open-source crawlers for LLMs: Crawl4AI vs Firecrawl vs Scrapy

Crawl4AI, self-hosted Firecrawl, Scrapy and Jina Reader compared on licence, browser rendering and output, plus the AI access question none of them answers.

Crawl4AI (Apache 2.0, Playwright, LLM-ready Markdown) is the most direct open-source crawler for LLM pipelines. Firecrawl can be self-hosted under AGPL-3.0 with the same API as its cloud. Scrapy (BSD-3) handles high-volume HTML without a browser. Jina Reader is free at r.jina.ai within rate limits, but its models are non-commercial. None shows whether AI crawlers can read your site; a free CrawlCheck scan does.

Share of all scans carrying each finding named abovePAGE_IS_MOSTLY_CODE29.2%CONTENT_NEEDS_JAVASCRIPT1.1%Share of all scans carrying eachfinding named abovePAGE_IS_MOSTLY_CODE29.2%CONTENT_NEEDS_JAVASCRIPT1.1%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Four open-source projects come up when teams want to feed web pages to an LLM without paying per page: Crawl4AI, Firecrawl’s self-hosted edition, Scrapy and Jina Reader. Their licences differ in ways that matter for commercial use, and only some run a browser. CrawlCheck is included for the one job it does that none of them do, not as an open-source alternative.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

Licences, browsers and output #

ProjectLicenceRuns a browserOutputNotes
Crawl4AIApache 2.0Yes, Playwright (Chromium)LLM-ready Markdown, structured extractionv0.9.4 (23 Sep 2026), about 84,600 GitHub stars
Firecrawl (self-hosted)AGPL-3.0; SDKs and some UI MITYesMarkdown, JSONDocker Compose; the cloud version has extra features; about 186,500 stars
ScrapyBSD-3-ClauseNo, not by itselfWhatever your spider yieldsAbout 64,500 stars; the high-volume HTML workhorse
Jina ReaderService code on GitHub; models CC-BY-NC 4.0 (not open source for commercial use)Hosted serviceMarkdown or JSONFree at r.jina.ai: 20 requests a minute without a key, 500 with a free key
CrawlCheckNot open source (hosted service)No; measures the HTML delivered before JavaScriptGraded report; signed per-domain answer by APIChecks whether AI crawlers can read a site; free scans

Sources: Crawl4AI, Firecrawl, Scrapy, Jina Reader.

Licences matter more than stars #

Apache 2.0 (Crawl4AI) and BSD-3 (Scrapy) let you build a commercial product on the code with few conditions. AGPL-3.0 (Firecrawl) requires you to publish your changes if you offer the modified software to users over a network. Jina’s models are licensed for non-commercial use only; commercial production needs a commercial licence. Read each licence before you build a product on it.

Choosing one #

The question none of them answers #

All four fetch pages for you. None tells you whether AI companies’ crawlers can fetch your pages, because they are not those crawlers and sites treat crawlers by name. Across CrawlCheck’s scans, content only exists after JavaScript runs on 1.1% of sites and the page is mostly code on 29.2%: an open-source crawler with a browser reads those pages fine while most AI crawlers do not. A free CrawlCheck scan shows what each AI crawler receives.

What each one gets from a page that needs JavaScript #

ProjectA page whose text is built by JavaScript
Crawl4AIReads the rendered text, because it drives Chromium through Playwright
Firecrawl (self-hosted)Reads the rendered text, because it renders pages in a browser
ScrapyGets the HTML the server sent, the same as most AI crawlers, unless you add a browser integration
Jina ReaderReturns its own rendered and cleaned Markdown
CrawlCheckMeasures the HTML the server sent, on purpose, because that is what most AI crawlers read

That difference cuts both ways. A browser-based crawler gets more out of a JavaScript-heavy site, which is what you want when you are collecting data. It also means its output is a poor guide to what an AI company’s crawler sees on that site. If a page reads perfectly in Crawl4AI and comes back nearly empty from Scrapy, most AI crawlers are getting the Scrapy version.

Crawling politely #

All four will fetch pages as fast as you let them. Running your own crawler means carrying the obligations the big crawlers carry:

Sites increasingly treat unknown crawlers as hostile, so a crawler that ignores these rules tends to end up behind challenges and 403s, which costs you the data you were after.

A quick test of what an AI crawler sees on your own site #

If you already run one of these tools, compare two fetches of the same page: one through a browser-based crawler and one as plain HTML. With Scrapy installed, scrapy fetch --nolog https://example.com/ | wc -c gives the size of the HTML the server sends; view the same page in a browser with JavaScript disabled to see what that HTML says. A large gap between that and the rendered page is the content most AI crawlers miss. A CrawlCheck scan makes the same comparison per crawler identity and adds what each was allowed by robots.txt.

Licence, browser and politeness settled, the choice usually comes down to scale: Scrapy for volume, a browser-based crawler for pages that need rendering, and a hosted reader for one-off pages.

Where CrawlCheck is weaker #

CrawlCheck is a hosted service, not open source: you cannot self-host or extend it like the projects above. If you need a crawler you control end to end, pick one of them.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is the best open-source crawler for LLMs?

Crawl4AI is the most direct fit: Apache 2.0, runs Playwright, outputs LLM-ready Markdown. Self-hosted Firecrawl offers the same API as its cloud under AGPL-3.0. Scrapy (BSD-3) suits high-volume HTML without a browser. Jina Reader is a free hosted service with rate limits.

Is Firecrawl open source?

Yes. The repository is mainly AGPL-3.0, with SDKs and some UI components under MIT, and it can be self-hosted with Docker Compose. The cloud version at firecrawl.dev has additional features.

Is Jina Reader free?

Basic use is free by prepending r.jina.ai to a URL: 20 requests a minute without an API key and 500 with a free key. The Reader models are licensed CC-BY-NC 4.0, so commercial production use needs a commercial licence.

Does Scrapy render JavaScript?

Not by itself. Scrapy fetches and parses HTML at high volume and is usually paired with a browser service when pages need JavaScript.

Can an open-source crawler tell me if ChatGPT can read my site?

No. It fetches pages as itself, not as GPTBot or ClaudeBot from those companies' networks. A CrawlCheck scan compares what each AI crawler identity is served, free.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works