Content Signals is a robots.txt directive, created by Cloudflare in September 2025, that declares what may be done with your content after a crawler has already downloaded it. It does not say who gets in — that is still Disallow's job. It says whether what was taken can be used to index, to ground a generative answer, or to train a model.
That distinction is the whole proposal. In the words of Cloudflare's own announcement, robots.txt "does not, however, let them know what they are able to do with your content after accessing it".
Our position, plainly: it is a statement of intent with claimed legal weight, not an access control. If you want a bot not to take your content, Content Signals will not do it. For that you need Disallow per user-agent, a WAF, or Bot Management. Adding it costs almost nothing; believing it does something today is the mistake.
The specification published at contentsignals.org defines exactly three signals:
search — building a search index and returning results: hyperlinks and short excerpts. The official definition explicitly excludes AI-generated summaries. search=yes does not authorize AI Overviews.ai-input — feeding the content into a model in real time: RAG, grounding, generative answers.ai-train — training or fine-tuning models.The only permitted values are yes and no.
There is a fourth, use, which Cloudflare introduced in July 2026 and describes literally as an optional extension under testing. Its values, from least to most permissive: use=immediate (interact, do not store or reuse), use=reference (index, cite and link — the default), use=full (summarize and reproduce).
A detail that matters: use is not documented on contentsignals.org. We checked today against the site's bundle: the word does not appear. The primary source for the spec and Cloudflare's production implementation are out of sync.
The canonical form from the official generator:
User-Agent: *
Content-Signal: ai-train=no, search=yes, ai-input=no
Allow: /
It goes inside a robots.txt group, alongside User-Agent, Allow and Disallow. Comma-separated key=value pairs.
You can segment by bot by repeating the block with a different User-Agent. You can also segment by path, with a not-so-obvious syntax: the path goes before the pairs, separated by a space and with no comma.
User-Agent: *
Content-Signal: /blog/ ai-train=no, search=yes, ai-input=no
Allow: /blog/
A warning about the grammar: there is no ABNF, no RFC, no versioned normative document. And it shows. Cloudflare writes the field three different ways depending on the source: Content-Signal: with a space after the comma on the blog and on contentsignals.org, Content-signal: with no spaces in its documentation, and Content-Signal: search=yes,ai-train=no,use=reference in actual production output. In practice this is irrelevant (robots.txt field names are treated as case-insensitive), but it is symptomatic: a spec whose author writes it three ways does not have a normative grammar, it has examples.
This is the nuance most often lost. The normative text says that if the site operator does not include a signal for a given use, it neither grants nor restricts permission with respect to that use. It is neutral, not negative.
That is why the default Cloudflare applied to its customers omits ai-input on purpose: it does not know the rights holder's preference and does not want to guess it.
Which opens the uncomfortable question of why it did guess the others. Cloudflare rolled the policy out automatically across 3.8 million domains with managed robots.txt, and in July added use=reference to them without any action by the rights holder. Under the policy's own legal theory, the absence of a signal is neutral but a signal that is present is a declaration of intent. Which means: millions of sites are declaring an intent that nobody on those sites ever formulated.
The official boilerplate says it in capitals: any restriction expressed via content signals constitutes an express reservation of rights under Article 4 of Directive (EU) 2019/790. That article conditions the text and data mining exception on rights not having been reserved "in an appropriate manner, such as machine-readable means".
It also tries to lean on a contract theory: the text opens by saying that, as a condition of accessing the site, whoever accesses it agrees to comply with these signals.
Now the fine print on the same site: Cloudflare warns that courts and regulators could conclude that robots.txt imposes no legally enforceable obligations, and recommends consulting a lawyer. The blog tones it down further still: it says the closing paragraph reminds readers that these signals might have legal effect in some jurisdictions.
Our reading: the assertion in capitals is a claim by an interested party, not a judicial determination. No court has yet ruled that a Content-Signal: constitutes a valid Art. 4(3) reservation.
The available case law cuts both ways. In Kneschke v. LAION, the OLG Hamburg (10 December 2025, ref. 5 U 104/24) reversed the lower court's generous reading: a reservation in natural language did not meet the machine-readability standard as of the time of use. In Content Signals' favor: it confirms the reservation must be machine-readable, and this one clearly is. Against it: the court anchors validity to the technology deployed at the moment of the act. Being machine-readable in theory is not the same as being read by the machines that matter.
And the scope is the EU only. There is no equivalent theory articulated for the United States, where the debate is fair use, nor for Argentina or Latin America.
Cloudflare admits it across all three of its sources. The blog: the signals express preferences, they are not technical countermeasures against scraping, and some companies may simply ignore them. The documentation is more direct still: if you want to force the block rather than request it, use AI Crawl Control.
On the crawler side, the evidence is worse than ambiguous:
google/robotstxt, adding content-signal to the kUnsupportedTags list. Google recognizes it so as not to report it as unknown. It does not act on it.There is also a competing standard with more process legitimacy: the IETF's aipref working group defines a Content-Usage field with a two-category vocabulary (train-ai, search) and y/n values, on its way to Proposed Standard. It shares not one token with Content Signals. A verifiable irony: contentsignals.org's own robots.txt declares that it uses vocabulary from the IETF standard and then emits Content-Signal, which is not from there.
Total adoption asymmetry: millions of emitters, zero confirmed receivers. As a technical control, its measured effectiveness today is zero.
Our recommendation, which is a judgment and not a certainty:
Content-Signal will not do it. Use Disallow per user-agent, a WAF, or AI Crawl Control. Nothing else.content-signal from its unsupported list, or the IETF draft advancing to RFC. A recommendation with no criterion for measurement is incomplete.One assumption we cannot verify: the absence of a public announcement does not prove that no crawler consumes it quietly. In Google's case, the commit does prove non-use.
Writing this line by hand on a Tiendanube store has a prior problem: the platform generates the robots.txt and does not let you edit it. The file has to respond at the root of the same domain and there is no way to upload it.
That is why the IndexNow Connect robots.txt editor has a per-signal selector — search, ai-input, ai-train and use — that builds the line for you, alongside the per-bot switches and the disallowed paths. It reads your real file, tells you who appears to be writing it, and lets you design a new one.
Two things it does not do, worth stating:
Allow rules, Crawl-delay, or per-bot groups with partial rules. What survives: the * group's signals, the disallowed paths, full per-bot blocks, and the sitemaps.The editor even warns you when you declare ai-train=no and leave AI bots unblocked, because the intuitive reading ("I already blocked it") is exactly the wrong one. If you want to check first which bots actually reach your site today, the AI bot accessibility test goes out over the network with each User-Agent and tells you.
Content Signals is the declarative layer. The layer that actually moves the needle today — getting your new URLs to the indexes looking for them, fast — is a different one, and it is what IndexNow Connect solves: notifying at the moment a product changes, instead of waiting for someone to come by. If you are starting from zero with the protocol, begin with what IndexNow is and continue with what GEO is to understand the other side of the problem.
No. It is a declaration of use, not an access control. It coexists with Allow: / in Cloudflare's default policy. The actual blocking is done by Disallow: / lines per user-agent (GPTBot, ClaudeBot, CCBot, Google-Extended and company), which are a separate and older mechanism.
search=yes, do I show up in AI Overviews?No — at least not according to the official definition. The spec text explicitly excludes AI-generated summaries from the scope of search. That is declared with ai-input. That said, no crawler is honoring either signal today, so in practice this determines nothing.
Nothing, and that is deliberate. The normative text says the absence of a signal neither grants nor restricts permission: it is neutral. It is not equivalent to a no. That is why Cloudflare omits ai-input in its default: it does not want to guess its customers' preference.
Google Search Console may report "Syntax not understood" for this and other new directives. Cloudflare states it has observed no impact on crawl rates or on SEO. The real risk of touching a robots.txt does not come from this line: it comes from a badly written Disallow rule, which really can drop you out of the search engines.
They are distinct and non-exclusive mechanisms: Content-Signal (Cloudflare) and Content-Usage (IETF) share neither vocabulary nor values. With cc-signals and llms.txt also in play, the most likely outcome is not that the best one wins, but that none reaches critical mass on the crawler side — which is the only side that matters. Emitting both lines costs two lines; assuming either is being obeyed does not.