CookingMetrics Data-Driven Business
Martín Garay·September 3, 2026·9 min readTechnical SEOGEO · AIGuides

The bot accessibility test: three layers, two questions

Almost everyone checks that their products are reachable by bots. Almost nobody checks that their checkout isn't.

Both halves matter equally. A bot that can't reach your catalog costs you visibility in search engines and in generated answers. A bot that does reach your cart, your account area, or your internal search pages costs you crawl budget — and sometimes something worse.

This article explains how you verify each side, and why the answers come from three independent layers that fail in different ways.

The three layers that keep a crawler out

When you say "this bot can't get in," you're conflating three mechanisms that have nothing to do with each other. Fixing each one falls to a different person.

Layer 1 — robots.txt. The site decides it, and compliant bots obey. It's voluntary: nothing stops a crawler from ignoring it. You resolve it by looking up the bot's product token inside the file, not its User-Agent. They're different fields: Screaming Frog uses the token screaming frog seo spider and goes out on the wire with the UA Screaming Frog SEO Spider/20.0.

Layer 2 — the HTTP response to the bot's User-Agent. The server, the WAF, or the CDN decides this. It's a real access cut, not a convention: it can block even when robots.txt allows, and it doesn't depend on anyone's goodwill. The statuses that count as a block are 401, 403, 407, 429, and 451.

Layer 3 — indexing directives. X-Robots-Tag in the headers and <meta name="robots"> in the <head>. They let the bot through but keep the page out. The bot accesses it, reads it, and the page still doesn't make it into the index.

All three can contradict each other. A permissive robots.txt with a WAF 403 on top of it means "blocked." A clean 200 with noindex in the header means "gets in but useless." That's why a traffic light isn't enough: you need to see the raw data from each layer separately.

Why the order matters

When you have to boil the three down to one verdict, this is the precedence:

  1. Network error — there was no response, so there's nothing to interpret.
  2. HTTP block (401/403/407/429/451) — beats robots.txt, because the bot never gets as far as reading the page.
  3. Disallow in robots.txt — what's forbidden, even if the server would have answered.
  4. Noindex — gets in, but isn't indexed.
  5. 5xx / 4xx / indeterminate robots — statuses that are neither yes nor no.

First what cuts off access, then what forbids it, and last what lets it in without indexing. A Disallow for GPTBot on a site that also returns 403 to it isn't redundant information: if you drop the WAF block tomorrow, that first rule is still standing.

The five reading errors that produce false negatives

Verifying this by hand, with curl and good intentions, fails in specific places. These are the ones that most often produce an "it's allowed" that isn't true.

That last point is worth being explicit about, because the intuitive reading is the wrong one: declaring ai-train=no and leaving AI bots unblocked does not block them. It's a reservation of rights, not a lock.

How we do it, and what it doesn't do

In IndexNow Connect the test runs from the dashboard, against up to 5 URLs and a catalog of 46 entries split across five groups: AI (21), search engines (10), SEO (6), social and messaging (8), and one browser row. Each row shows a column per layer plus the verdict.

Three decisions worth spelling out, because they have visible consequences:

There's a browser row, and it's the most important one. Chrome on macOS, with no robots token. Without it you have nothing to compare against: if the browser also gets a 403, the problem isn't the bot, it's the URL. It's the first question to ask about any block.

Three tokens get no request. Google-Extended, Applebot-Extended, and the legacy anthropic-ai aren't crawlers: they're AI-use switches over what Googlebot and Applebot have already fetched. Sending them a probe would mean inventing a bot that doesn't exist. For those, we report the robots.txt verdict and nothing else.

The ceiling is real and it's declared. Maximum 40 HTTP probes per run, 4 in parallel, a 6-second timeout, 200 KB of body read per response. With two URLs and the full catalog, 86 probes are requested and 40 run: 46 go untested, the screen says so with a number, and those rows come back as "untested" in yellow, never as allowed. A silent limit reads as "I tested everything," and that's worse than not measuring at all.

What it does not do, plainly stated:

The test lives alongside the rest of the dashboard — take a look at IndexNow Connect if you want to see how it fits with automatically submitting URLs to IndexNow when you publish or edit a product.

What to check on each half

What should get in. Home, main categories, a representative product page, the sitemap. Check the search engine group and the AI group separately: it's common for Googlebot to pass clean while GPTBot gets a 403 from a WAF rule nobody remembers putting there. If the browser gets in and the bot doesn't, the block is either deliberate or accidental, but it exists.

What should not get in. Checkout, cart, account area, internal search pages with a query string, faceted filters, API endpoints. Here the expected result is disallow, and an allow is the finding. Test with the query string included: robots.txt evaluation takes pathname + search, and a rule written for /search doesn't necessarily cover /search?q=sneakers.

Two details that change the result on the second half:

If your robots.txt is written by your platform or rewritten for you by Cloudflare, chances are neither of these is the way you think it is. It's also worth reading what Content Signals in robots.txt are and how to keep your site accessible to AI bots with Cloudflare.

Frequently asked questions

Does a Disallow in robots.txt guarantee the bot won't get in?

No. The Robots Exclusion Protocol is voluntary: it describes how a crawler that wants to comply should interpret the file, it doesn't impose anything. If you need a real access cut, the layer is HTTP: authentication, WAF, or CDN rules. That's why the test measures both separately.

If I don't have a robots.txt, is everything allowed?

Yes for access, but you lose the standard sitemap declaration. Watch the nuance: a 4xx means "there's no file" and therefore everything is allowed, while a 5xx does not amount to allowed. It's an unknown state, and treating it as permission means inventing an answer the server never gave.

Do noindex and Disallow do the same thing?

No, and combining them is usually a mistake. If you forbid the URL in robots.txt, the bot can't read the noindex you put inside it. Google documents this in its noindex guide: for the directive to be honored, the page has to be crawlable. Forbidding in robots.txt keeps them from spending crawl budget; noindex keeps the page out of the index.

Why does checkout need to be blocked if it doesn't rank anyway?

For two different reasons. The first is crawl budget: every cart URL with unique parameters is a new URL the bot discovers and requests, and that competes with your product pages. The second is that account pages and purchase flows sometimes expose data or states that were never meant to leave the site. The second is rarer and more expensive.

How many bots is it worth checking?

Fewer than you think, but the right ones. Start with the AI group and the browser row, which is the default. The browser tells you whether the URL works; the AI group is where the blocks nobody consciously decided show up, because many WAFs ship with "AI scrapers" rules enabled out of the box. The SEO and social groups add noise unless you're diagnosing a specific link-preview problem or a tool that can't crawl your site.