If you run a Tiendanube store, you can't edit your robots.txt. The platform generates it and gives you no access to the file. And if you're on Cloudflare with managed robots.txt, you can't either: Cloudflare rewrites the whole thing from its # BEGIN Cloudflare Managed content block, and whatever you edit there gets overwritten on the next regeneration.
That's the starting point for the IndexNow Connect robots.txt editor. It doesn't assume you write the file. It assumes someone else writes it for you, and works from there.
What the screen does is concrete: it reads your real file, tells you who appears to be writing it, lets you design a new one with one toggle per bot, and — if you want — publishes it through a Cloudflare Worker that you paste into your own account.
Every time you open the screen, a live request goes out to https://yourdomain/robots.txt. With a 6-second timeout, a 500 KB cap, and an anti-SSRF guard. A 4xx is read as "there's no file"; a 5xx stays unknown, because a server error is not the same as "everything allowed".
The analysis runs over that text. It gives you back:
cdn.shopify.com is enough to flag Shopify. That's why the screen says "who appears to be writing it" and, when nothing matches, "we couldn't identify it"./. The question the editor answers is "does this bot get in or not?", not the verdict for a specific URL.AdsBot-Google unnamed while disallowed paths exist, tokens nobody recognizes (possible typo), duplicate tokens, no Sitemap:, Crawl-delay present, and one specific note: you declared ai-train=no but there are AI bots left unblocked.That last note exists because the intuitive reading is the wrong one. Content-Signal declares permitted use, it does not block access. If you want the detail on that line — what each signal means, who honors it and who doesn't — it's in content-signals-robots-txt.
The grid has 45 toggles across four groups: AI (21), search engines (10), SEO (6) and social (8). Each one switched on means "this bot gets in".
When you switch one off, the generated file doesn't add a line to the * group. It creates that bot its own group:
User-agent: GPTBot
Disallow: /
That's the only way the protocol has to exclude one and let the rest in. And it has a consequence worth understanding: since a bot's own group beats the wildcard, that bot stops reading the * rules. That's REP by design, not a decision of ours. Cloudflare does exactly the same in its block.
Now, the persistence decision. In the database we store the blocked ones, not the allowed ones.
The reason is the default. When we add a new bot to the catalog tomorrow, with "blocked" that bot starts out allowed — which is what you expect — whereas with "allowed" it would end up disallowed without anyone deciding it. A change we make to the catalog can't block traffic you never asked to block. There's a dedicated test for that.
The on-screen editor works with the set of allowed bots, because that's what each toggle marks. The conversion between the two models lives at the edge, between the screen and the database.
Beyond the toggles, the editor handles:
search, ai-input, ai-train) plus Cloudflare's experimental use. If you declare none, the block isn't written: with no signals the file stays clean.* group, one per line.Sitemap: lines, which go at the end.User-agent: AdsBot-Google group is generated with the same paths repeated. Google's AdsBot does not obey the * group: if you don't name it, it keeps coming in. It's the same solution Tiendanube uses.The generated file does not preserve everything the original had. Four things survive: the * group's signals, the * group's Disallow paths, full per-bot blocks, and the Sitemap: lines.
What's lost:
Allow: rules. A fixed Allow: / is always written.Crawl-delay lines. (Googlebot ignores them anyway; Bing and Yandex honor them.)Googlebot-Image: Disallow: /photos/ becomes either "fully blocked" or nothing.There's also a real bug worth naming. The generator writes the bot's name, not its token. For 44 of the 45 they match once lowercased. The exception is Screaming Frog: the catalog declares the token screaming frog seo spider, but the file comes out with User-agent: Screaming Frog. Switching that toggle off doesn't produce a rule the crawler recognizes as its own. We verified it by generating a file with every bot blocked and analyzing it again: it's the only one that comes back as not blocked.
We don't validate the syntax of what you write, either. Paths aren't checked for starting with /, sitemaps aren't checked for being valid URLs or for pointing at your domain. Any non-empty line goes in as-is. That's why the screen warns you, right above the button: a badly built robots.txt can take your site out of the search engines.
Here's the most important limitation, and the one people forget most.
"Save and publish" does two things: it saves the configuration and it flips a flag. With that flag on, /robotsfile/yourdomain starts returning the file. Your domain keeps serving exactly what it always served.
The file goes live only once the Cloudflare Worker has the route. The Worker is a snippet you paste into your own Cloudflare account — there's no integration with the Cloudflare API, you set up the route by hand — and it intercepts exactly two URLs: the IndexNow key (/{hex}.txt) and /robots.txt. Anything else passes straight through.
And it fails open, by design. If our API doesn't return 200, the Worker does the original fetch and your domain serves its usual robots.txt. An outage on our side can't leave anyone's store without a robots.txt — or with an empty one.
Why a Worker is needed: it's the only mechanism we have to answer a URL at the root of someone else's domain. The platform doesn't allow uploading files and the protocol requires the same host. There's no shortcut through another domain. It's exactly the same problem we solve for the IndexNow key with Cloudflare Workers, and the Cloudflare technical guide for SEO and GEO covers the rest of the pieces.
The screen distinguishes three situations that people confuse and that have different fixes:
The third state isn't declared just because a button was pressed. It's verified: the real file's text is compared against the generated one, in full, exact except for leading and trailing whitespace. If they don't match, we don't say it's published.
Two honest caveats about that check. First: the comparison is boolean, there's no visual diff. Second: the "generated" file it compares against is the draft you have on screen. If you flip toggles and hit Generate without publishing, the card may say "your domain serves something else" even though the published version really is live.
And the underlying limitation: the check only happens when someone opens the screen. There's no scheduled verification and no alerts. Nothing re-verifies later if the Worker stops serving it. There's no history or versioning either: we store the last configuration, the publication date and the update date.
The endpoint the Worker consumes regenerates the file on every request from the configuration; we don't store the file, we store the design. It responds with Cache-Control: no-store, because unpublishing has to be visible on the next request and not in five minutes. And it returns 404 when nothing is published: that 404 is part of the design, it's what makes the Worker pass through. Unpublishing flips the flag off and your domain instantly goes back to its original robots.txt, without touching Cloudflare. The configuration is kept: unpublishing turns off the light, it doesn't throw away the design.
If on top of deciding who gets in you want to measure what happens when they do, IndexNow Connect connects your store, notifies the search engines every time you change a product, and shows you whether AI bots are actually arriving. The robots.txt editor is one piece of that, not the whole product.
You can design the file, analyze it and download it. You can't publish it on your domain. On Tiendanube the file is served by the platform, and the Cloudflare Worker is the only way we have to answer a URL at the root of your domain. If your site runs on a platform where you can upload the file yourself, download it and upload it.
It stops coming in, if it honors the REP. Disallow is a preference, not a technical barrier: a crawler can ignore it. Real blocking takes WAF or Bot Management rules, which is what Cloudflare's own docs recommend. And don't confuse blocking with Content-Signal: the signal declares permitted use, it doesn't prevent access.
Because the generator builds the file from the editor's configuration, it doesn't patch it. It keeps * group signals, * group Disallow paths, full per-bot blocks and the Sitemap: lines. Everything else — specific Allow: rules, Crawl-delay, groups with partial rules — doesn't survive. Before publishing, look at the generated file on screen and compare it with the raw one we show you above.
Nothing on your site. The Worker queries our API and only responds if a 200 comes back. Otherwise it falls through to the original fetch and your domain serves its usual robots.txt. It's an explicit decision in the Worker's design: it fails open.
The editor tells you what your file declares; it doesn't prove what happens on the wire. That's what the accessibility test for AI bots is for — it goes out with each bot's User-Agent and checks the real response. They're two deliberately different tools: one works on the file, the other on the traffic.