CookingMetrics Data-Driven Business
Martín Garay·September 3, 2026·10 min readGEO · AICloudflareTechnical SEO

Content Signals in robots.txt: what it declares, how to write it, and why it blocks nothing

Content Signals is a robots.txt directive, created by Cloudflare in September 2025, that declares what may be done with your content after a crawler has already downloaded it. It does not say who gets in — that is still Disallow's job. It says whether what was taken can be used to index, to ground a generative answer, or to train a model.

That distinction is the whole proposal. In the words of Cloudflare's own announcement, robots.txt "does not, however, let them know what they are able to do with your content after accessing it".

Our position, plainly: it is a statement of intent with claimed legal weight, not an access control. If you want a bot not to take your content, Content Signals will not do it. For that you need Disallow per user-agent, a WAF, or Bot Management. Adding it costs almost nothing; believing it does something today is the mistake.

The three signals (and the fourth, which is experimental)

The specification published at contentsignals.org defines exactly three signals:

The only permitted values are yes and no.

There is a fourth, use, which Cloudflare introduced in July 2026 and describes literally as an optional extension under testing. Its values, from least to most permissive: use=immediate (interact, do not store or reuse), use=reference (index, cite and link — the default), use=full (summarize and reproduce).

A detail that matters: use is not documented on contentsignals.org. We checked today against the site's bundle: the word does not appear. The primary source for the spec and Cloudflare's production implementation are out of sync.

How it is written

The canonical form from the official generator:

User-Agent: *
Content-Signal: ai-train=no, search=yes, ai-input=no
Allow: /

It goes inside a robots.txt group, alongside User-Agent, Allow and Disallow. Comma-separated key=value pairs.

You can segment by bot by repeating the block with a different User-Agent. You can also segment by path, with a not-so-obvious syntax: the path goes before the pairs, separated by a space and with no comma.

User-Agent: *
Content-Signal: /blog/ ai-train=no, search=yes, ai-input=no
Allow: /blog/

A warning about the grammar: there is no ABNF, no RFC, no versioned normative document. And it shows. Cloudflare writes the field three different ways depending on the source: Content-Signal: with a space after the comma on the blog and on contentsignals.org, Content-signal: with no spaces in its documentation, and Content-Signal: search=yes,ai-train=no,use=reference in actual production output. In practice this is irrelevant (robots.txt field names are treated as case-insensitive), but it is symptomatic: a spec whose author writes it three ways does not have a normative grammar, it has examples.

The absence of a signal is not a "no"

This is the nuance most often lost. The normative text says that if the site operator does not include a signal for a given use, it neither grants nor restricts permission with respect to that use. It is neutral, not negative.

That is why the default Cloudflare applied to its customers omits ai-input on purpose: it does not know the rights holder's preference and does not want to guess it.

Which opens the uncomfortable question of why it did guess the others. Cloudflare rolled the policy out automatically across 3.8 million domains with managed robots.txt, and in July added use=reference to them without any action by the rights holder. Under the policy's own legal theory, the absence of a signal is neutral but a signal that is present is a declaration of intent. Which means: millions of sites are declaring an intent that nobody on those sites ever formulated.

The legal status: a self-serving assertion

The official boilerplate says it in capitals: any restriction expressed via content signals constitutes an express reservation of rights under Article 4 of Directive (EU) 2019/790. That article conditions the text and data mining exception on rights not having been reserved "in an appropriate manner, such as machine-readable means".

It also tries to lean on a contract theory: the text opens by saying that, as a condition of accessing the site, whoever accesses it agrees to comply with these signals.

Now the fine print on the same site: Cloudflare warns that courts and regulators could conclude that robots.txt imposes no legally enforceable obligations, and recommends consulting a lawyer. The blog tones it down further still: it says the closing paragraph reminds readers that these signals might have legal effect in some jurisdictions.

Our reading: the assertion in capitals is a claim by an interested party, not a judicial determination. No court has yet ruled that a Content-Signal: constitutes a valid Art. 4(3) reservation.

The available case law cuts both ways. In Kneschke v. LAION, the OLG Hamburg (10 December 2025, ref. 5 U 104/24) reversed the lower court's generous reading: a reservation in natural language did not meet the machine-readability standard as of the time of use. In Content Signals' favor: it confirms the reservation must be machine-readable, and this one clearly is. Against it: the court anchors validity to the technology deployed at the moment of the act. Being machine-readable in theory is not the same as being read by the machines that matter.

And the scope is the EU only. There is no equivalent theory articulated for the United States, where the debate is fair use, nor for Argentina or Latin America.

What it does not do: the hard evidence

Cloudflare admits it across all three of its sources. The blog: the signals express preferences, they are not technical countermeasures against scraping, and some companies may simply ignore them. The documentation is more direct still: if you want to force the block rather than request it, use AI Crawl Control.

On the crawler side, the evidence is worse than ambiguous:

There is also a competing standard with more process legitimacy: the IETF's aipref working group defines a Content-Usage field with a two-category vocabulary (train-ai, search) and y/n values, on its way to Proposed Standard. It shares not one token with Content Signals. A verifiable irony: contentsignals.org's own robots.txt declares that it uses vocabulary from the IETF standard and then emits Content-Signal, which is not from there.

Total adoption asymmetry: millions of emitters, zero confirmed receivers. As a technical control, its measured effectiveness today is zero.

So, do I add it or not?

Our recommendation, which is a judgment and not a certainty:

  1. If the goal is to stop them taking your content: Content-Signal will not do it. Use Disallow per user-agent, a WAF, or AI Crawl Control. Nothing else.
  2. If the goal is to build evidence of express reservation with a view to future EU litigation: the marginal cost is close to zero and the option value is not nil. Add it.
  3. How to know if this starts to matter: there are two objective, cheap signals to monitor — some relevant operator removing content-signal from its unsupported list, or the IETF draft advancing to RFC. A recommendation with no criterion for measurement is incomplete.

One assumption we cannot verify: the absence of a public announcement does not prove that no crawler consumes it quietly. In Google's case, the commit does prove non-use.

How we generate it

Writing this line by hand on a Tiendanube store has a prior problem: the platform generates the robots.txt and does not let you edit it. The file has to respond at the root of the same domain and there is no way to upload it.

That is why the IndexNow Connect robots.txt editor has a per-signal selector — search, ai-input, ai-train and use — that builds the line for you, alongside the per-bot switches and the disallowed paths. It reads your real file, tells you who appears to be writing it, and lets you design a new one.

Two things it does not do, worth stating:

The editor even warns you when you declare ai-train=no and leave AI bots unblocked, because the intuitive reading ("I already blocked it") is exactly the wrong one. If you want to check first which bots actually reach your site today, the AI bot accessibility test goes out over the network with each User-Agent and tells you.

Content Signals is the declarative layer. The layer that actually moves the needle today — getting your new URLs to the indexes looking for them, fast — is a different one, and it is what IndexNow Connect solves: notifying at the moment a product changes, instead of waiting for someone to come by. If you are starting from zero with the protocol, begin with what IndexNow is and continue with what GEO is to understand the other side of the problem.

Frequently asked questions

Does Content Signals block ChatGPT or Claude?

No. It is a declaration of use, not an access control. It coexists with Allow: / in Cloudflare's default policy. The actual blocking is done by Disallow: / lines per user-agent (GPTBot, ClaudeBot, CCBot, Google-Extended and company), which are a separate and older mechanism.

If I set search=yes, do I show up in AI Overviews?

No — at least not according to the official definition. The spec text explicitly excludes AI-generated summaries from the scope of search. That is declared with ai-input. That said, no crawler is honoring either signal today, so in practice this determines nothing.

What happens if I do not declare a signal?

Nothing, and that is deliberate. The normative text says the absence of a signal neither grants nor restricts permission: it is neutral. It is not equivalent to a no. That is why Cloudflare omits ai-input in its default: it does not want to guess its customers' preference.

Does adding this line to my robots.txt break anything?

Google Search Console may report "Syntax not understood" for this and other new directives. Cloudflare states it has observed no impact on crawl rates or on SEO. The real risk of touching a robots.txt does not come from this line: it comes from a badly written Disallow rule, which really can drop you out of the search engines.

Is it worth waiting for the IETF standard?

They are distinct and non-exclusive mechanisms: Content-Signal (Cloudflare) and Content-Usage (IETF) share neither vocabulary nor values. With cc-signals and llms.txt also in play, the most likely outcome is not that the best one wins, but that none reaches critical mass on the crawler side — which is the only side that matters. Emitting both lines costs two lines; assuming either is being obeyed does not.