Guides · 2026-10-03 · By VSNARY | Emmanuel Orta · 0 views
Open-source crawlers for LLMs: Crawl4AI vs Firecrawl vs Scrapy
Crawl4AI, self-hosted Firecrawl, Scrapy and Jina Reader compared on licence, browser rendering and output, plus the AI access question none of them answers.
Crawl4AI (Apache 2.0, Playwright, LLM-ready Markdown) is the most direct open-source crawler for LLM pipelines. Firecrawl can be self-hosted under AGPL-3.0 with the same API as its cloud. Scrapy (BSD-3) handles high-volume HTML without a browser. Jina Reader is free at r.jina.ai within rate limits, but its models are non-commercial. None shows whether AI crawlers can read your site; a free CrawlCheck scan does.
Four open-source projects come up when teams want to feed web pages to an LLM without paying per page: Crawl4AI, Firecrawl’s self-hosted edition, Scrapy and Jina Reader. Their licences differ in ways that matter for commercial use, and only some run a browser. CrawlCheck is included for the one job it does that none of them do, not as an open-source alternative.
Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.
Licences, browsers and output #
| Project | Licence | Runs a browser | Output | Notes |
|---|---|---|---|---|
| Crawl4AI | Apache 2.0 | Yes, Playwright (Chromium) | LLM-ready Markdown, structured extraction | v0.9.4 (23 Sep 2026), about 84,600 GitHub stars |
| Firecrawl (self-hosted) | AGPL-3.0; SDKs and some UI MIT | Yes | Markdown, JSON | Docker Compose; the cloud version has extra features; about 186,500 stars |
| Scrapy | BSD-3-Clause | No, not by itself | Whatever your spider yields | About 64,500 stars; the high-volume HTML workhorse |
| Jina Reader | Service code on GitHub; models CC-BY-NC 4.0 (not open source for commercial use) | Hosted service | Markdown or JSON | Free at r.jina.ai: 20 requests a minute without a key, 500 with a free key |
| CrawlCheck | Not open source (hosted service) | No; measures the HTML delivered before JavaScript | Graded report; signed per-domain answer by API | Checks whether AI crawlers can read a site; free scans |
Sources: Crawl4AI, Firecrawl, Scrapy, Jina Reader.
Licences matter more than stars #
Apache 2.0 (Crawl4AI) and BSD-3 (Scrapy) let you build a commercial product on the code with few conditions. AGPL-3.0 (Firecrawl) requires you to publish your changes if you offer the modified software to users over a network. Jina’s models are licensed for non-commercial use only; commercial production needs a commercial licence. Read each licence before you build a product on it.
Choosing one #
- LLM-ready Markdown, self-hosted, permissive licence: Crawl4AI.
- Same API locally as the Firecrawl cloud: self-hosted Firecrawl, with the AGPL in mind.
- Millions of plain HTML pages: Scrapy.
- A quick Markdown copy of one page, no setup: Jina Reader within its rate limits.
The question none of them answers #
All four fetch pages for you. None tells you whether AI companies’ crawlers can fetch your pages, because they are not those crawlers and sites treat crawlers by name. Across CrawlCheck’s scans, content only exists after JavaScript runs on 1.1% of sites and the page is mostly code on 29.2%: an open-source crawler with a browser reads those pages fine while most AI crawlers do not. A free CrawlCheck scan shows what each AI crawler receives.
What each one gets from a page that needs JavaScript #
| Project | A page whose text is built by JavaScript |
|---|---|
| Crawl4AI | Reads the rendered text, because it drives Chromium through Playwright |
| Firecrawl (self-hosted) | Reads the rendered text, because it renders pages in a browser |
| Scrapy | Gets the HTML the server sent, the same as most AI crawlers, unless you add a browser integration |
| Jina Reader | Returns its own rendered and cleaned Markdown |
| CrawlCheck | Measures the HTML the server sent, on purpose, because that is what most AI crawlers read |
That difference cuts both ways. A browser-based crawler gets more out of a JavaScript-heavy site, which is what you want when you are collecting data. It also means its output is a poor guide to what an AI company’s crawler sees on that site. If a page reads perfectly in Crawl4AI and comes back nearly empty from Scrapy, most AI crawlers are getting the Scrapy version.
Crawling politely #
All four will fetch pages as fast as you let them. Running your own crawler means carrying the obligations the big crawlers carry:
- Honour robots.txt. Scrapy’s project template turns this on with
ROBOTSTXT_OBEY = True; check the equivalent setting in whichever tool you use rather than assuming it is on. - Name yourself. Send a user-agent that identifies your crawler and links to a page explaining it, so a site owner can contact you or write a rule for you instead of blocking a whole network.
- Limit concurrency per host. A small site behind shared hosting can be slowed by a few parallel browser sessions.
- Cache what you fetched. Refetching unchanged pages costs the site and you.
Sites increasingly treat unknown crawlers as hostile, so a crawler that ignores these rules tends to end up behind challenges and 403s, which costs you the data you were after.
A quick test of what an AI crawler sees on your own site #
If you already run one of these tools, compare two fetches of the same page: one through a browser-based crawler and one as plain HTML. With Scrapy installed, scrapy fetch --nolog https://example.com/ | wc -c gives the size of the HTML the server sends; view the same page in a browser with JavaScript disabled to see what that HTML says. A large gap between that and the rendered page is the content most AI crawlers miss. A CrawlCheck scan makes the same comparison per crawler identity and adds what each was allowed by robots.txt.
Licence, browser and politeness settled, the choice usually comes down to scale: Scrapy for volume, a browser-based crawler for pages that need rendering, and a hosted reader for one-off pages.
Where CrawlCheck is weaker #
CrawlCheck is a hosted service, not open source: you cannot self-host or extend it like the projects above. If you need a crawler you control end to end, pick one of them.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What is the best open-source crawler for LLMs?
Crawl4AI is the most direct fit: Apache 2.0, runs Playwright, outputs LLM-ready Markdown. Self-hosted Firecrawl offers the same API as its cloud under AGPL-3.0. Scrapy (BSD-3) suits high-volume HTML without a browser. Jina Reader is a free hosted service with rate limits.
Is Firecrawl open source?
Yes. The repository is mainly AGPL-3.0, with SDKs and some UI components under MIT, and it can be self-hosted with Docker Compose. The cloud version at firecrawl.dev has additional features.
Is Jina Reader free?
Basic use is free by prepending r.jina.ai to a URL: 20 requests a minute without an API key and 500 with a free key. The Reader models are licensed CC-BY-NC 4.0, so commercial production use needs a commercial licence.
Does Scrapy render JavaScript?
Not by itself. Scrapy fetches and parses HTML at high volume and is usually paired with a browser service when pages need JavaScript.
Can an open-source crawler tell me if ChatGPT can read my site?
No. It fetches pages as itself, not as GPTBot or ClaudeBot from those companies' networks. A CrawlCheck scan compares what each AI crawler identity is served, free.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.