CrawlCheck

Guides · 2026-10-02 · By · 0 views

Markdown for agents: serving text/markdown on request

An agent that sends Accept: text/markdown can be given the page's words without the scripts, styles and navigation. How content negotiation works for this, the headers that must come back, the caching trap, and what an intermediary may add on your behalf.

Markdown for agents is HTTP content negotiation applied to AI clients: when a request carries Accept: text/markdown, the server returns the same page as Markdown with Content-Type: text/markdown; charset=utf-8 and Vary: Accept, at the same URL. Browsers keep receiving HTML. The Vary header is essential, or an edge cache can serve one representation to the other kind of client. Cloudflare's Markdown for Agents converts HTML at the edge on Pro, Business and Enterprise plans, adds x-markdown-tokens, and sets a content-signal header allowing search, ai-input and ai-train when the origin has not set one.

A language model reading a web page does not need the navigation, the cookie banner, the inline styles or the script bundle. It needs the words and their structure: headings, lists, tables, links. Converting a page to that form is work every agent repeats on every fetch, and often does badly. Content negotiation lets the site do it once, correctly, and hand the result to any client that asks. This guide covers how to serve Markdown at the same URL, the response headers that matter, the cache configuration that makes it safe, and what happens when an intermediary does the conversion for you.

One URL, two representations, chosen by the Accept headerSAME URLGET /guides/exampleone addressAccept: text/htmla browserAccept: text/markdownan agentHTML pagenav, scripts, stylesMarkdownthe words, structure keptVary: Accept keeps the two cached apartWithout Vary: Accept, an edge cache can hand the Markdown to browsers or the HTML to agents.

How the negotiation works #

HTTP has always let a client say which formats it prefers through the Accept request header. A browser sends something like text/html,application/xhtml+xml,…. An agent that wants Markdown sends Accept: text/markdown. The server reads the header and returns the matching representation of the same resource at the same URL. There is no separate address to discover, no .md suffix to guess, and nothing changes for browsers or search crawlers, which never ask for Markdown.

$ curl -sI https://crawlcheck.io/guides/xml-sitemap-audit -H 'Accept: text/markdown'
content-type: text/markdown; charset=utf-8
vary: Accept

The response headers that must come back #

HeaderValueWhy
Content-Typetext/markdown; charset=utf-8the client confirms it got Markdown, not HTML that happened to match
Varyincludes Acceptcaches store the HTML and Markdown versions separately
Link (optional)rel="canonical" to the HTML URLkeeps one canonical address for the resource
x-markdown-tokens (Cloudflare)estimated token countlets an agent budget before it reads
content-signal (Cloudflare default)ai-train=yes, search=yes, ai-input=yes if the origin set nonea permission statement you may not have chosen

The caching trap #

The failure mode is a shared cache that ignores the Accept header. The first client to request a URL decides what is cached; if it was an agent, browsers get raw Markdown until the entry expires, and if it was a browser, agents get HTML forever. Vary: Accept tells caches to key on the header. Some CDNs ignore Vary for cache keys unless configured, so test from outside: request the URL with and without the header and confirm each returns its own type, then repeat after the cache has warmed. Stale and split machine-file caches are a separate but related problem covered in the machine-file caching guide.

Doing it at the edge #

Cloudflare's Markdown for Agents feature performs the conversion at the edge for HTML responses up to 2 MB, on Pro, Business and Enterprise plans, and sets Vary: Accept itself. It also adds a content-signal header permitting search, AI input and AI training when the origin has not sent one. That default may not match what your robots.txt says, so if you have chosen a different policy, send your own content-signal header from the origin; the meaning of each value is in the content signals guide. Doing the conversion yourself gives control over what is kept: you can drop the footer and related-posts block, keep tables as Markdown tables, and put the page's answer first.

What to put in the Markdown #

Keep everything that carries meaning: the title, the date, the author, headings in order, lists, tables, code, and links with absolute URLs. Drop navigation, repeated footers, share buttons, cookie notices and anything that only exists to make the HTML page work. If the HTML page has structured data, keep the facts in the text; the Markdown representation has nowhere to put JSON-LD. Make sure the words match the HTML. A Markdown version that says something the HTML page does not is a different document at the same address, and an agent comparing them will trust neither.

Implementing it at the origin #

At the origin the logic is small. Read the Accept header; if it names text/markdown with a quality value at least as high as text/html, render the Markdown representation, otherwise render HTML as usual. Generate the Markdown from the same source as the HTML, the article's stored content, rather than by converting the rendered HTML back, so the two cannot drift. Return Vary: Accept on both representations, not only the Markdown one, because a cache that stored the HTML without a Vary header will never look for a second variant.

Decide what happens to requests for text/markdown on pages that have no Markdown version, such as a checkout or a search results page. Returning the HTML with its normal content type is correct under HTTP and tells the client honestly what it got; returning 406 Not Acceptable is also correct but breaks clients that list Markdown first without a fallback. Most sites should return HTML.

Testing it #

Request the same URL twice from outside, once with Accept: text/markdown and once with a browser's Accept header, and compare the content types and the Vary header. Repeat both requests in reverse order after the cache has warmed, to prove the cache keeps them apart. Then read the Markdown: headings in order, tables intact, links absolute, no leftover navigation, and the same facts as the HTML page.

Does it help visibility? #

It reduces what an agent has to fetch and parse, and removes a conversion step where text is often lost or mangled, particularly tables and code. It does not make a page rank or get cited by itself; no engine has said it prefers sites that negotiate Markdown. Treat it like an llms.txt file: a low-cost courtesy to machine readers that makes the text you already have easier to use. The pages that need it most are the ones that are mostly code, where the Markdown can be a small fraction of the HTML.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

How does an AI agent request Markdown instead of HTML?

By sending Accept: text/markdown. A server that supports it returns the same page as Markdown at the same URL.

What response headers should a Markdown response have?

Content-Type: text/markdown; charset=utf-8 and Vary: Accept. Optionally a canonical Link header pointing to the HTML URL.

Why is Vary: Accept required?

Without it, a shared cache can store whichever version was requested first and serve it to everyone, so browsers get Markdown or agents get HTML.

Does Cloudflare convert pages to Markdown automatically?

Its Markdown for Agents feature converts HTML responses up to 2 MB at the edge for clients that send Accept: text/markdown, on Pro, Business and Enterprise plans.

Does Cloudflare's conversion change my AI permissions?

It adds a content-signal header allowing search, ai-input and ai-train when the origin has not set one. Send your own header if your policy differs.

Does serving Markdown improve rankings or citations?

No engine has said so. It makes the text easier and cheaper for agents to use, which is a delivery improvement, not a ranking signal.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works