CrawlCheck

Guides · 2026-10-03 · By · 0 views

Best llms.txt generators and validators compared

Firecrawl, Yoast, Mintlify and CrawlCheck compared on generating llms.txt, keeping it current and checking it actually reaches AI crawlers.

Yoast SEO generates and refreshes llms.txt automatically for WordPress and Shopify, free. Mintlify hosts llms.txt and llms-full.txt for documentation sites. Firecrawl's free generator builds llms.txt and llms-full.txt for any site, and CrawlCheck drafts one from your declared pages. After publishing, CrawlCheck's free audit checks that AI crawlers actually receive the file as text, which no generator does.

Share of all scans carrying each finding named aboveNO_LLMS_TXT30.9%MACHINE_FILE_CACHE_STALE0.6%Share of all scans carrying eachfinding named aboveNO_LLMS_TXT30.9%MACHINE_FILE_CACHE_STALE0.6%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

An llms.txt file is a Markdown file at the root of a site that points AI tools to its most useful pages. The proposal at llmstxt.org (by Jeremy Howard) requires one thing, an H1 with the site’s name, followed by an optional summary blockquote and sections of links. Generating one is now easy; several tools do it automatically. The harder part is making sure the file you publish is actually served to AI crawlers, as text, with links that resolve.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

Generators and validators compared #

ToolWhat it doesKeeps it updated?Checks it is served to crawlers?Price
CrawlCheckDrafts an llms.txt from your homepage and declared pages; audits the live fileNo (draft only)Yes: checks /llms.txt on apex and www as AI crawlers receive it: status, content type, cache and whether cited pages resolveFree
Firecrawl llms.txt generatorCrawls a site and generates llms.txt and llms-full.txt (full page content)No; regenerate on demandNoFree; a free API key removes limits
Yoast SEOGenerates llms.txt for WordPress and Shopify sites, automatic or manual modeYes, automaticallyNoIncluded free in Yoast
MintlifyHosts llms.txt and llms-full.txt for documentation sitesYes, automaticallyNoPart of Mintlify docs hosting

Sources: Firecrawl, Yoast, Mintlify, CrawlCheck free tools and llms.txt audit.

Why generating the file is not the same as publishing it #

A file that exists in your CMS can still fail to reach a crawler. The patterns we measure: the URL answers 200 with an HTML page or a login wall instead of text; a cache keeps serving an old copy after you change the file; the apex and www hosts serve different files; or a bot rule refuses AI crawlers the file a browser receives. We documented one in an llms.txt that answered 200 with a login wall. Across all scans, 30.9% of sites have no llms.txt at all and 0.6% serve a stale cached machine file.

Which to use #

What a minimal valid llms.txt looks like #

Whichever tool writes the draft, the result should follow the order the proposal sets out: an H1 with the site’s name (the only required part), a blockquote summary, optional notes, then sections headed with H2 that list links. A section named “Optional” is, by convention, the links an AI tool can skip when it needs a shorter context.

# Example Fence Co

> Residential and commercial fence installation and repair in Denver, Colorado.

Family-run since 2011. Prices on the pricing page are current.

## Services

- [Wood fences](https://example.com/wood-fences/): materials, styles, typical cost
- [Fence repair](https://example.com/fence-repair/): what we repair and how fast

## Optional

- [Company history](https://example.com/about/)

Every link should point to a page that answers 200 for a crawler, carries the facts the note promises, and is not a redirect to the homepage. A generator that copies your navigation menu tends to list everything; a short file that names the ten pages that actually answer customers’ questions is more useful to a model with a limited context.

Five checks after you publish #

  1. Status 200, not a redirect to a page. An HTML page at /llms.txt is a soft 404 even when it answers 200.
  2. Content type text/plain or text/markdown. Some servers send text/html or force a download.
  3. The same file on apex and www. If example.com/llms.txt and www.example.com/llms.txt differ, a crawler that starts on the other host reads the other file.
  4. A fresh copy, not a cached one. After editing, request the file with a cache-busting query and without one; if they differ, a cache is serving the old file.
  5. Served to AI crawlers, not only to browsers. A bot rule that refuses GPTBot on the whole site refuses it on /llms.txt too.

You can run the first three with one command: curl -sI https://example.com/llms.txt and the same for the www host, reading the status line and the content-type header. The last two need a fetch as each crawler, which is what the free audit does.

llms.txt next to the other machine files #

FileTells a machineStandard
robots.txtWhat each crawler may fetchYes, RFC 9309
sitemap.xmlEvery URL you want discovered, with datesYes, sitemaps.org
llms.txtThe few pages worth reading first, with notesA proposal
llms-full.txtThe full text of those pages in one fileA convention used by some tools

llms.txt does not replace either standard file. A page refused in robots.txt stays refused even if llms.txt lists it, and a crawler that never reads llms.txt still discovers pages through the sitemap.

Where CrawlCheck is weaker #

CrawlCheck’s drafter works from your homepage and declared pages and does not keep the file updated; Yoast and Mintlify do that automatically, and Firecrawl also produces llms-full.txt with full page content, which CrawlCheck does not. CrawlCheck’s strength is the check after publishing.

Does llms.txt actually help? #

It is a proposal, not a standard the major AI crawlers have committed to read. Treat it as low-cost and low-certainty: worth publishing correctly, not worth publishing broken. The evidence so far is in does llms.txt actually work?, and the format in how to write an llms.txt and verify it.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is the best llms.txt generator?

For WordPress or Shopify sites on Yoast SEO, Yoast's built-in llms.txt is free and updates automatically. Mintlify documentation sites get llms.txt and llms-full.txt automatically. For any other site, Firecrawl's free generator and CrawlCheck's free drafter create a starting file. Whatever you use, confirm the live file is served to AI crawlers as text.

Is there a free llms.txt validator?

Yes. CrawlCheck's free llms.txt audit checks /llms.txt on the apex and www hosts as AI crawlers receive it: status, content type, cache freshness and whether the pages it cites resolve. A file that answers 200 with HTML or a login wall is not a working llms.txt.

What is the difference between llms.txt and llms-full.txt?

llms.txt is a short Markdown index with links to a site's key pages. llms-full.txt, produced by tools such as Firecrawl and Mintlify, includes the full content of each page in one file.

Does Yoast create llms.txt?

Yes. Yoast's llms.txt feature is free in Yoast SEO, Premium, WooCommerce SEO, AI+ and Yoast SEO for Shopify, and generates and refreshes the file automatically, with a manual mode for choosing content.

Do AI crawlers read llms.txt?

It is a proposal rather than a standard the major AI crawlers have committed to. Publishing it correctly is cheap; publishing it broken can serve crawlers an HTML page or login wall at that address.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works