CookingMetrics Data-Driven Business
Martín Garay·September 3, 2026·9 min readTechnical SEOCloudflareGuides

Our robots.txt editor: one toggle per bot, analysis of the real file, and publishing via Cloudflare

If you run a Tiendanube store, you can't edit your robots.txt. The platform generates it and gives you no access to the file. And if you're on Cloudflare with managed robots.txt, you can't either: Cloudflare rewrites the whole thing from its # BEGIN Cloudflare Managed content block, and whatever you edit there gets overwritten on the next regeneration.

That's the starting point for the IndexNow Connect robots.txt editor. It doesn't assume you write the file. It assumes someone else writes it for you, and works from there.

What the screen does is concrete: it reads your real file, tells you who appears to be writing it, lets you design a new one with one toggle per bot, and — if you want — publishes it through a Cloudflare Worker that you paste into your own account.

First it reads the file that's live right now

Every time you open the screen, a live request goes out to https://yourdomain/robots.txt. With a 6-second timeout, a 500 KB cap, and an anti-SSRF guard. A 4xx is read as "there's no file"; a 5xx stays unknown, because a server error is not the same as "everything allowed".

The analysis runs over that text. It gives you back:

That last note exists because the intuitive reading is the wrong one. Content-Signal declares permitted use, it does not block access. If you want the detail on that line — what each signal means, who honors it and who doesn't — it's in content-signals-robots-txt.

One toggle per bot, and the blocked ones are what we store

The grid has 45 toggles across four groups: AI (21), search engines (10), SEO (6) and social (8). Each one switched on means "this bot gets in".

When you switch one off, the generated file doesn't add a line to the * group. It creates that bot its own group:

User-agent: GPTBot
Disallow: /

That's the only way the protocol has to exclude one and let the rest in. And it has a consequence worth understanding: since a bot's own group beats the wildcard, that bot stops reading the * rules. That's REP by design, not a decision of ours. Cloudflare does exactly the same in its block.

Now, the persistence decision. In the database we store the blocked ones, not the allowed ones.

The reason is the default. When we add a new bot to the catalog tomorrow, with "blocked" that bot starts out allowed — which is what you expect — whereas with "allowed" it would end up disallowed without anyone deciding it. A change we make to the catalog can't block traffic you never asked to block. There's a dedicated test for that.

The on-screen editor works with the set of allowed bots, because that's what each toggle marks. The conversion between the two models lives at the edge, between the screen and the database.

What else goes into the generated file

Beyond the toggles, the editor handles:

What the generator loses (and it's worth saying out loud)

The generated file does not preserve everything the original had. Four things survive: the * group's signals, the * group's Disallow paths, full per-bot blocks, and the Sitemap: lines.

What's lost:

There's also a real bug worth naming. The generator writes the bot's name, not its token. For 44 of the 45 they match once lowercased. The exception is Screaming Frog: the catalog declares the token screaming frog seo spider, but the file comes out with User-agent: Screaming Frog. Switching that toggle off doesn't produce a rule the crawler recognizes as its own. We verified it by generating a file with every bot blocked and analyzing it again: it's the only one that comes back as not blocked.

We don't validate the syntax of what you write, either. Paths aren't checked for starting with /, sitemaps aren't checked for being valid URLs or for pointing at your domain. Any non-empty line goes in as-is. That's why the screen warns you, right above the button: a badly built robots.txt can take your site out of the search engines.

Publishing doesn't touch your domain

Here's the most important limitation, and the one people forget most.

"Save and publish" does two things: it saves the configuration and it flips a flag. With that flag on, /robotsfile/yourdomain starts returning the file. Your domain keeps serving exactly what it always served.

The file goes live only once the Cloudflare Worker has the route. The Worker is a snippet you paste into your own Cloudflare account — there's no integration with the Cloudflare API, you set up the route by hand — and it intercepts exactly two URLs: the IndexNow key (/{hex}.txt) and /robots.txt. Anything else passes straight through.

And it fails open, by design. If our API doesn't return 200, the Worker does the original fetch and your domain serves its usual robots.txt. An outage on our side can't leave anyone's store without a robots.txt — or with an empty one.

Why a Worker is needed: it's the only mechanism we have to answer a URL at the root of someone else's domain. The platform doesn't allow uploading files and the protocol requires the same host. There's no shortcut through another domain. It's exactly the same problem we solve for the IndexNow key with Cloudflare Workers, and the Cloudflare technical guide for SEO and GEO covers the rest of the pieces.

Three states, not two

The screen distinguishes three situations that people confuse and that have different fixes:

  1. Not published. Nothing is saved with the flag on.
  2. Published, but your domain still serves something else. The classic mistake: the route is missing from the Worker.
  3. Live. The real file matches the generated one.

The third state isn't declared just because a button was pressed. It's verified: the real file's text is compared against the generated one, in full, exact except for leading and trailing whitespace. If they don't match, we don't say it's published.

Two honest caveats about that check. First: the comparison is boolean, there's no visual diff. Second: the "generated" file it compares against is the draft you have on screen. If you flip toggles and hit Generate without publishing, the card may say "your domain serves something else" even though the published version really is live.

And the underlying limitation: the check only happens when someone opens the screen. There's no scheduled verification and no alerts. Nothing re-verifies later if the Worker stops serving it. There's no history or versioning either: we store the last configuration, the publication date and the update date.

The endpoint the Worker consumes regenerates the file on every request from the configuration; we don't store the file, we store the design. It responds with Cache-Control: no-store, because unpublishing has to be visible on the next request and not in five minutes. And it returns 404 when nothing is published: that 404 is part of the design, it's what makes the Worker pass through. Unpublishing flips the flag off and your domain instantly goes back to its original robots.txt, without touching Cloudflare. The configuration is kept: unpublishing turns off the light, it doesn't throw away the design.

If on top of deciding who gets in you want to measure what happens when they do, IndexNow Connect connects your store, notifies the search engines every time you change a product, and shows you whether AI bots are actually arriving. The robots.txt editor is one piece of that, not the whole product.

Frequently asked questions

Can I use the editor without Cloudflare?

You can design the file, analyze it and download it. You can't publish it on your domain. On Tiendanube the file is served by the platform, and the Cloudflare Worker is the only way we have to answer a URL at the root of your domain. If your site runs on a platform where you can upload the file yourself, download it and upload it.

If I block an AI bot, does it stop using my content?

It stops coming in, if it honors the REP. Disallow is a preference, not a technical barrier: a crawler can ignore it. Real blocking takes WAF or Bot Management rules, which is what Cloudflare's own docs recommend. And don't confuse blocking with Content-Signal: the signal declares permitted use, it doesn't prevent access.

Why does the generated file lose my current rules?

Because the generator builds the file from the editor's configuration, it doesn't patch it. It keeps * group signals, * group Disallow paths, full per-bot blocks and the Sitemap: lines. Everything else — specific Allow: rules, Crawl-delay, groups with partial rules — doesn't survive. Before publishing, look at the generated file on screen and compare it with the raw one we show you above.

What happens if your server goes down?

Nothing on your site. The Worker queries our API and only responds if a 200 comes back. Otherwise it falls through to the original fetch and your domain serves its usual robots.txt. It's an explicit decision in the Worker's design: it fails open.

How do I know whether the bots are actually getting in?

The editor tells you what your file declares; it doesn't prove what happens on the wire. That's what the accessibility test for AI bots is for — it goes out with each bot's User-Agent and checks the real response. They're two deliberately different tools: one works on the file, the other on the traffic.