CrawlCheck

Guides · 2026-10-02 · By · 0 views

Content Signals in robots.txt: search, ai-input, ai-train

Content-Signal is a robots.txt line that states what you allow done with your content after it is fetched: search, AI input, AI training. What each value means under Cloudflare's published policy, why it does not block anything, and the one combination that costs you citations.

Content-Signal is a robots.txt line, introduced in Cloudflare's Content Signals Policy in September 2025, that states preferences for three uses of content: search (indexing and showing links and short excerpts), ai-input (feeding content into AI models for grounded or real-time answers) and ai-train (training or fine-tuning models), each set to yes or no. An absent value neither grants nor refuses. A signal is not an access rule; crawling is still governed by User-agent and Disallow lines. CrawlCheck raises AI_OPTOUT_SET only when search or ai-input is refused, because ai-train=no costs no reach or citation.

Share of all scans carrying each finding named aboveAI_OPTOUT_SET5%Share of all scans carrying eachfinding named aboveAI_OPTOUT_SET5%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Robots.txt was designed to answer one question: may this crawler fetch this path. It has no way to say "you may fetch this for search, but not to train a model". Content signals add that second dimension. They are a single line in robots.txt stating which uses of the content the site permits once it has been fetched. This guide covers the three values, what each one actually affects, how they interact with ordinary crawl rules, and why one setting quietly removes a site from AI answers.

The three content signals and what refusing each one costsCONTENT-SIGNAL VALUECOVERSIF SET TO NOsearchindex + links + short excerptsdropped from classic resultsai-inputgrounding, RAG, live AI answersnot quoted in AI answersai-traintraining or fine-tuning modelsno effect on reach or citationA signal is a stated preference, not an access rule. The User-agent rules decide what can be fetched.

The syntax and the three values #

The line sits in robots.txt, typically inside the User-agent: * group, and lists comma-separated pairs:

User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /

Cloudflare's policy, published on 24 September 2025, defines the three values. search covers building a search index and returning links and short excerpts, and explicitly excludes AI-generated search summaries. ai-input covers putting content into a model at answer time: retrieval, grounding, and real-time use in generative answers. ai-train covers training or fine-tuning a model. Each takes yes or no. A value left out is not a default yes or a default no; the policy says the operator neither grants nor restricts that use.

A signal is a preference, not a wall #

Nothing in a content signal stops a request. A crawler that ignores the line can fetch every page the User-agent rules allow. That has two consequences. First, if you need a use prevented rather than discouraged, the crawl rules have to do it: disallow the training crawler by name. Second, the signal and the rules should agree. A file that says ai-train=no while giving GPTBot an unconditional Allow: / states one policy and enforces another. Our own robots.txt carries a comment to that effect: every value reads yes because everything is fetchable, and flipping ai-train to no would require disallowing the training crawlers as well.

Which setting costs you visibility #

Signal setEffect on being foundEffect on being quoted in AI answersCrawlCheck
no line at allnonenoneno finding
ai-train=no onlynonenoneno finding; noted in the engines section
ai-input=nononeremoved, for engines that honour itAI_OPTOUT_SET
search=noremoved from classic results, for engines that honour itusually removed tooAI_OPTOUT_SET
signal says no, rules allow everythingdepends on the crawlerdepends on the crawlerthe rules are what is enforced

The distinction that matters is between training and answering. Refusing training is a coherent position that costs nothing CrawlCheck measures: an engine can still reach, read and quote the site. Refusing ai-input asks answer engines not to use the page when composing an answer, which is the mechanism by which a page gets cited. We originally scored any opt-out as a failure, including ai-train=no, and corrected it, because punishing a deliberate and harmless choice is a measurement error. Today AI_OPTOUT_SET fires only on ai-input=no or search=no, on 5% of scans.

Who reads the signal #

The signal is a published policy with a reservation-of-rights statement, not a protocol with mandatory compliance. Whether a given operator honours it is that operator's decision, and most have not said publicly. Treat it as a clear statement of intent that may carry legal weight in some jurisdictions, and use crawl rules for anything you need enforced. Note also that intermediaries can set a signal on your behalf: Cloudflare's Markdown for Agents conversion adds a content-signal header reading yes for all three uses when the origin has not set one, as covered in serving Markdown to agents.

Choosing a setting #

If the goal is to be cited by answer engines, leave search and ai-input at yes. Set ai-train according to your view on model training; it does not affect citation. If you refuse training, back it with a Disallow: / group for the crawlers whose operators state they are used for training, such as GPTBot and CCBot, plus Google's Google-Extended token, and make sure those groups do not drop your other rules, as explained in how named robots.txt groups work. The difference between a training crawler and an answer-time fetcher for each operator is laid out in eight crawlers, eight companies.

Signals, robots rules and meta tags together #

Content signals sit beside two older mechanisms, and the three answer different questions. Robots.txt User-agent and Disallow lines decide whether a path may be fetched at all. Page-level directives such as noindex, nosnippet or max-snippet, in a meta robots tag or an X-Robots-Tag header, decide what a search engine may do with a page it has fetched, and Google applies nosnippet to its AI features as well as to classic snippets. Content signals state site-wide preferences for three categories of use. When they conflict, the most restrictive mechanism a given crawler honours is the one that applies to it.

That means a site does not need content signals to keep a page out of AI summaries on Google; nosnippet already does that, at the cost of snippets in search. Signals are most useful for operators who have no other way to say "answer with this, do not train on it" in one place. The page-level side is covered in meta robots and X-Robots-Tag.

Common mistakes #

Four mistakes recur. Putting the line outside any group, before the first User-agent line, where a parser that reads lines only within groups may ignore it. Spelling a value ai_train or aitrain, which matches nothing. Setting ai-input=no while meaning to refuse only training, which removes the site from AI answers it wanted to appear in. And setting a signal in robots.txt while a CDN or plugin sends a different content-signal header on every page. Read both before deciding what your site says.

Checking what your site says #

Fetch /robots.txt as a crawler would, not from a browser tab that may show a cached copy, and read the Content-Signal line and the groups around it. Confirm the values are the ones you chose, that the line sits in the group you intended, and that the crawl rules enforce any refusal you care about. The robots.txt checker reports the signal next to each AI crawler's effective rules.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is Content-Signal in robots.txt?

A line stating which uses of a site's content are permitted after it is fetched: search, ai-input and ai-train, each yes or no. It comes from Cloudflare's Content Signals Policy, published in September 2025.

What is the difference between ai-input and ai-train?

ai-input covers using content at answer time, such as retrieval and grounding for AI answers. ai-train covers training or fine-tuning a model.

Does Content-Signal block crawlers?

No. It is a stated preference. Whether a page can be fetched is still decided by the User-agent, Allow and Disallow rules.

Will ai-train=no stop my site being cited by ChatGPT or Perplexity?

No. Refusing training does not affect reaching, reading or quoting a page. Refusing ai-input or search is what can remove a site from AI answers.

What does it mean if a signal is missing?

Under the policy, an absent value neither grants nor restricts that use.

How do I enforce a refusal to train?

Disallow the training crawlers by name in robots.txt, for example GPTBot, CCBot and Google-Extended, and repeat your other rules inside those groups.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works