Almost everyone checks that their products are reachable by bots. Almost nobody checks that their checkout isn't.
Both halves matter equally. A bot that can't reach your catalog costs you visibility in search engines and in generated answers. A bot that does reach your cart, your account area, or your internal search pages costs you crawl budget — and sometimes something worse.
This article explains how you verify each side, and why the answers come from three independent layers that fail in different ways.
When you say "this bot can't get in," you're conflating three mechanisms that have nothing to do with each other. Fixing each one falls to a different person.
Layer 1 — robots.txt. The site decides it, and compliant bots obey. It's voluntary: nothing stops a crawler from ignoring it. You resolve it by looking up the bot's product token inside the file, not its User-Agent. They're different fields: Screaming Frog uses the token screaming frog seo spider and goes out on the wire with the UA Screaming Frog SEO Spider/20.0.
Layer 2 — the HTTP response to the bot's User-Agent. The server, the WAF, or the CDN decides this. It's a real access cut, not a convention: it can block even when robots.txt allows, and it doesn't depend on anyone's goodwill. The statuses that count as a block are 401, 403, 407, 429, and 451.
Layer 3 — indexing directives. X-Robots-Tag in the headers and <meta name="robots"> in the <head>. They let the bot through but keep the page out. The bot accesses it, reads it, and the page still doesn't make it into the index.
All three can contradict each other. A permissive robots.txt with a WAF 403 on top of it means "blocked." A clean 200 with noindex in the header means "gets in but useless." That's why a traffic light isn't enough: you need to see the raw data from each layer separately.
When you have to boil the three down to one verdict, this is the precedence:
First what cuts off access, then what forbids it, and last what lets it in without indexing. A Disallow for GPTBot on a site that also returns 403 to it isn't redundant information: if you drop the WAF block tomorrow, that first rule is still standing.
Verifying this by hand, with curl and good intentions, fails in specific places. These are the ones that most often produce an "it's allowed" that isn't true.
/old that redirects to a forbidden /private is not accessible. And if the redirect crosses hosts, the file that governs is the other host's.User-agent: Google does not cover Googlebot. The protocol requires an exact match (RFC 9309, section 2.2.1). Matching it "just in case" means reporting a prohibition the site never wrote.User-agent lines wrong. Two or more in a row share the rules that follow. Splitting them in the wrong place applies a Disallow to the wrong bot.X-Robots-Tag as plain text. The header supports per-bot scoping: X-Robots-Tag: googlebot: noindex. Searching the raw string for the word "noindex" flags GPTBot and bingbot over a rule written for Googlebot. You have to separate the scope prefix from directives that take values, like max-snippet: -1.Allow and still be barred from training. They're different axes.That last point is worth being explicit about, because the intuitive reading is the wrong one: declaring ai-train=no and leaving AI bots unblocked does not block them. It's a reservation of rights, not a lock.
In IndexNow Connect the test runs from the dashboard, against up to 5 URLs and a catalog of 46 entries split across five groups: AI (21), search engines (10), SEO (6), social and messaging (8), and one browser row. Each row shows a column per layer plus the verdict.
Three decisions worth spelling out, because they have visible consequences:
There's a browser row, and it's the most important one. Chrome on macOS, with no robots token. Without it you have nothing to compare against: if the browser also gets a 403, the problem isn't the bot, it's the URL. It's the first question to ask about any block.
Three tokens get no request. Google-Extended, Applebot-Extended, and the legacy anthropic-ai aren't crawlers: they're AI-use switches over what Googlebot and Applebot have already fetched. Sending them a probe would mean inventing a bot that doesn't exist. For those, we report the robots.txt verdict and nothing else.
The ceiling is real and it's declared. Maximum 40 HTTP probes per run, 4 in parallel, a 6-second timeout, 200 KB of body read per response. With two URLs and the full catalog, 86 probes are requested and 40 run: 46 go untested, the screen says so with a number, and those rows come back as "untested" in yellow, never as allowed. A silent limit reads as "I tested everything," and that's worse than not measuring at all.
What it does not do, plainly stated:
User-agent: Screaming Frog instead of its real token, so the crawler doesn't recognize that rule as its own.The test lives alongside the rest of the dashboard — take a look at IndexNow Connect if you want to see how it fits with automatically submitting URLs to IndexNow when you publish or edit a product.
What should get in. Home, main categories, a representative product page, the sitemap. Check the search engine group and the AI group separately: it's common for Googlebot to pass clean while GPTBot gets a 403 from a WAF rule nobody remembers putting there. If the browser gets in and the bot doesn't, the block is either deliberate or accidental, but it exists.
What should not get in. Checkout, cart, account area, internal search pages with a query string, faceted filters, API endpoints. Here the expected result is disallow, and an allow is the finding. Test with the query string included: robots.txt evaluation takes pathname + search, and a rule written for /search doesn't necessarily cover /search?q=sneakers.
Two details that change the result on the second half:
* group. If you don't name it explicitly, it keeps hitting the paths you forbade for everyone. You have to repeat the rules for it in its own group.* rules. It's the only way the protocol has to exclude one bot and let the rest in, but it means that bot stops seeing all your general Disallow lines. That's an effect of the protocol, not of the tool.If your robots.txt is written by your platform or rewritten for you by Cloudflare, chances are neither of these is the way you think it is. It's also worth reading what Content Signals in robots.txt are and how to keep your site accessible to AI bots with Cloudflare.
Disallow in robots.txt guarantee the bot won't get in?No. The Robots Exclusion Protocol is voluntary: it describes how a crawler that wants to comply should interpret the file, it doesn't impose anything. If you need a real access cut, the layer is HTTP: authentication, WAF, or CDN rules. That's why the test measures both separately.
Yes for access, but you lose the standard sitemap declaration. Watch the nuance: a 4xx means "there's no file" and therefore everything is allowed, while a 5xx does not amount to allowed. It's an unknown state, and treating it as permission means inventing an answer the server never gave.
noindex and Disallow do the same thing?No, and combining them is usually a mistake. If you forbid the URL in robots.txt, the bot can't read the noindex you put inside it. Google documents this in its noindex guide: for the directive to be honored, the page has to be crawlable. Forbidding in robots.txt keeps them from spending crawl budget; noindex keeps the page out of the index.
For two different reasons. The first is crawl budget: every cart URL with unique parameters is a new URL the bot discovers and requests, and that competes with your product pages. The second is that account pages and purchase flows sometimes expose data or states that were never meant to leave the site. The second is rarer and more expensive.
Fewer than you think, but the right ones. Start with the AI group and the browser row, which is the default. The browser tells you whether the URL works; the AI group is where the blocks nobody consciously decided show up, because many WAFs ship with "AI scrapers" rules enabled out of the box. The SEO and social groups add noise unless you're diagnosing a specific link-preview problem or a tool that can't crawl your site.