Cloudflare sits in front of one in five websites (approximate figure, according to searchVIU). That makes it the gatekeeper of crawling: every WAF rule, every challenge and every cache policy decides what Googlebot, Bingbot and AI bots see before a single byte reaches your server. That's where what searchVIU calls the GEO paradox lives: blocking all AI bots protects your content, but erases you from the generative answers where purchases are decided today. GEO starts at your CDN, not at your content.
A guide for engineers: features, fields and bots by exact name, with their trade-offs. Looking for the jargon-free version? Start with the step-by-step implementation guide.
Only proxied records (orange cloud) pass through your Cloudflare zone: on DNS only (gray cloud) no Workers, WAF, cache or Crawler Hints run. First check: apex and www in orange.
A real case that cost us hours of debugging: SaaS stores whose apex had A records pointing straight at the platform's IPs. With that, Cloudflare routes the root domain to the platform's infrastructure and your zone gets bypassed — Workers included — even though the proxy shows as active. The fix: replace those A records with a proxied CNAME @ → {store}.{platform}.com, same as www; the full diagnosis is in the IndexNow key with Workers case.
Of the SSL/TLS encryption modes, only one is defensible: Full (strict), with origin certificate validation.
With Cloudflare's free origin certificate, Full (strict) costs nothing: there's no real trade-off here.
The short rule: aggressive for static assets, HTML only with invalidation. searchVIU dedicates a section to it ("Caching Affects What Crawlers See"): crawlers can receive outdated content from the edge, so purge the cache when you publish changes.
For ecommerce it's worse: cached HTML with an old price or stock level is what Bing indexes — and what ChatGPT and Copilot answer with afterwards. The foundation is still a sitemap with an honest lastmod: if the cache lies and the sitemap does too, the crawler is left with no reliable signals.
Bot Fight Mode (Free plan) challenges automated traffic with no possible exceptions: no Skip rules, no allowlist. Super Bot Fight Mode (Pro and up) adds the Definitely automated / Verified bots toggles and, crucially, Skip rules.
Cloudflare validates verified bots (Googlebot, Bingbot and company) by IP and behavior, and exposes them in cf.client.bot, available on all plans (cf.bot_management.verified_bot is Enterprise-only). The docs say verified bots are exempt from challenges, but there are false positives reported in the community: verify instead of trusting.
The classic symptom: 403s to Googlebot or Bingbot, discovered when Search Console or Bing Webmaster Tools can't read the sitemap. Four-step diagnosis:
Googlebot, bingbot) and look for Block or Challenge events.cf.client.bot and put it first in the order: an earlier Block wins.In August 2026, Search Engine Journal reported a case — an r/TechSEO thread, without confirmation from Cloudflare — of AI bot blocking that kept Googlebot from indexing because of its "Search + Training" classification. The direction of the risk: anti-AI rules can trample search crawling.
"AI bots" are not one single thing. They split by purpose:
| Crawler | Purpose | robots.txt |
|---|---|---|
| GPTBot | Training (OpenAI) | Respects it |
| OAI-SearchBot | ChatGPT search | Respects it |
| ChatGPT-User | User action in ChatGPT | According to OpenAI, it "may not apply" |
| ClaudeBot | Training (Anthropic) | Respects it |
| Claude-SearchBot | Claude search | Respects it |
| Claude-User | User action | Anthropic states it respects it |
| PerplexityBot | Perplexity search (doesn't train) | Respects it |
| Perplexity-User | User action | "Generally ignores it" |
| Google-Extended | Control token: training + Gemini grounding | Token in robots.txt, no UA of its own |
Two clarifications almost nobody makes:
For stores: allow the search bots — OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Bingbot and Googlebot — and make a conscious decision about the training ones. The search bots generate the citations (and sales) in the answers.
Key notice: since July 1, 2025 ("Content Independence Day"), Cloudflare blocks AI crawlers by default on new zones (the exact scope depends on onboarding). If your zone is newer than that, review AI Crawl Control — all plans, with Allow, Block and Charge actions — and explicitly allow the search bots. Charge comes from pay-per-crawl (HTTP 402), still in private beta.
Last piece: the Content Signals Policy (Sep 24, 2025) adds signals to the managed robots.txt — Content-Signal: search=yes, ai-train=no — already on more than 3.8 million domains. But these are signals, not enforcement. The wall is Robotcop (December 2024): network-level robots.txt enforcement against the bot that declares one thing and does another. A signal for the one that complies; a wall for the one that doesn't.
Cloudflare is already an IndexNow submitter. Crawler Hints (Cache → Configuration, free on all plans) notifies search engines when it detects changed content, using the IndexNow protocol. It requires proxied records and excludes URLs that respond 4xx or worse.
The nuance: it fires on cache MISS, a passive heuristic: Cloudflare infers the change from cache state, without knowing what changed or whether it matters. A per-event push is a different thing: the "product updated" webhook carries the exact URL at the exact moment, including removals, which a MISS never represents well. They don't compete: turn on Crawler Hints and add per-event push on top. For stores, that's IndexNow Connect: it listens to your catalog's webhooks (additions, changes, removals) and sends each URL to IndexNow instantly, with no heuristics and without touching your store's code.
A Worker runs before your origin. Two concrete SEO uses:
The free plan includes 100,000 daily requests: for these uses, it's more than enough.
If you do the SEO but don't manage the client's Cloudflare, don't ask for admin: ask for read-only access to Security → Events, Security → Bots and AI Crawl Control. Enough to audit without touching anything.
www proxied; no A records pointing at a SaaS platform's IPs.cf.client.bot first in the order; Security → Events with no blocks on Googlebot/Bingbot.It shouldn't: the docs say verified bots are exempt, but there are false positives reported in the community. On the Free plan there are no exceptions or Skip rules, so if you see 403s to Googlebot in Security → Events, your options are turning off Bot Fight Mode or moving to Pro and using Super Bot Fight Mode with a Skip rule.
Both, and your invalidation decides which. It improves the response time users and crawlers see, but stale HTML served from the edge pushes outdated prices and stock into the index and into AI answers. If you can't purge on publish, don't cache HTML.
It's a business decision about training, not about visibility. Blocking GPTBot doesn't remove you from ChatGPT search — that's governed by OAI-SearchBot, which you should allow. The serious mistake is blocking both by reflex.
Because the 403 isn't emitted by your server: Cloudflare emits it before the request reaches the origin. A curl from your network won't reproduce it because you're not Googlebot. Filter Security → Events by user-agent, identify the rule and create a Skip rule first in the order.
Connect your store for free → — Crawler Hints notifies when the cache guesses; IndexNow Connect notifies when your catalog actually changes.