CookingMetrics Data-Driven Business
Martín Garay·August 19, 2026·8 min readCloudflareTechnical SEO

Cloudflare for SEO and GEO: a technical configuration guide

Cloudflare sits in front of one in five websites (approximate figure, according to searchVIU). That makes it the gatekeeper of crawling: every WAF rule, every challenge and every cache policy decides what Googlebot, Bingbot and AI bots see before a single byte reaches your server. That's where what searchVIU calls the GEO paradox lives: blocking all AI bots protects your content, but erases you from the generative answers where purchases are decided today. GEO starts at your CDN, not at your content.

A guide for engineers: features, fields and bots by exact name, with their trade-offs. Looking for the jargon-free version? Start with the step-by-step implementation guide.

DNS: proxied or nothing

Only proxied records (orange cloud) pass through your Cloudflare zone: on DNS only (gray cloud) no Workers, WAF, cache or Crawler Hints run. First check: apex and www in orange.

A real case that cost us hours of debugging: SaaS stores whose apex had A records pointing straight at the platform's IPs. With that, Cloudflare routes the root domain to the platform's infrastructure and your zone gets bypassed — Workers included — even though the proxy shows as active. The fix: replace those A records with a proxied CNAME @ → {store}.{platform}.com, same as www; the full diagnosis is in the IndexNow key with Workers case.

TLS: Full (strict), always

Of the SSL/TLS encryption modes, only one is defensible: Full (strict), with origin certificate validation.

With Cloudflare's free origin certificate, Full (strict) costs nothing: there's no real trade-off here.

Cache: speed yes, stale HTML no

The short rule: aggressive for static assets, HTML only with invalidation. searchVIU dedicates a section to it ("Caching Affects What Crawlers See"): crawlers can receive outdated content from the edge, so purge the cache when you publish changes.

For ecommerce it's worse: cached HTML with an old price or stock level is what Bing indexes — and what ChatGPT and Copilot answer with afterwards. The foundation is still a sitemap with an honest lastmod: if the cache lies and the sitemap does too, the crawler is left with no reliable signals.

Bots: where indexing breaks

Bot Fight Mode (Free plan) challenges automated traffic with no possible exceptions: no Skip rules, no allowlist. Super Bot Fight Mode (Pro and up) adds the Definitely automated / Verified bots toggles and, crucially, Skip rules.

Cloudflare validates verified bots (Googlebot, Bingbot and company) by IP and behavior, and exposes them in cf.client.bot, available on all plans (cf.bot_management.verified_bot is Enterprise-only). The docs say verified bots are exempt from challenges, but there are false positives reported in the community: verify instead of trusting.

The classic symptom: 403s to Googlebot or Bingbot, discovered when Search Console or Bing Webmaster Tools can't read the sitemap. Four-step diagnosis:

  1. Security → Events: filter by user-agent (Googlebot, bingbot) and look for Block or Challenge events.
  2. Identify the rule ID and the service that fired it (custom WAF, Super Bot Fight Mode, rate limiting).
  3. Create a Skip rule with cf.client.bot and put it first in the order: an earlier Block wins.
  4. Re-verify with URL Inspection (Search Console) and its equivalent in Bing Webmaster Tools.

In August 2026, Search Engine Journal reported a case — an r/TechSEO thread, without confirmation from Cloudflare — of AI bot blocking that kept Googlebot from indexing because of its "Search + Training" classification. The direction of the risk: anti-AI rules can trample search crawling.

AI crawlers: block by purpose, not by brand

"AI bots" are not one single thing. They split by purpose:

Crawler Purpose robots.txt
GPTBot Training (OpenAI) Respects it
OAI-SearchBot ChatGPT search Respects it
ChatGPT-User User action in ChatGPT According to OpenAI, it "may not apply"
ClaudeBot Training (Anthropic) Respects it
Claude-SearchBot Claude search Respects it
Claude-User User action Anthropic states it respects it
PerplexityBot Perplexity search (doesn't train) Respects it
Perplexity-User User action "Generally ignores it"
Google-Extended Control token: training + Gemini grounding Token in robots.txt, no UA of its own

Two clarifications almost nobody makes:

For stores: allow the search bots — OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Bingbot and Googlebot — and make a conscious decision about the training ones. The search bots generate the citations (and sales) in the answers.

Key notice: since July 1, 2025 ("Content Independence Day"), Cloudflare blocks AI crawlers by default on new zones (the exact scope depends on onboarding). If your zone is newer than that, review AI Crawl Control — all plans, with Allow, Block and Charge actions — and explicitly allow the search bots. Charge comes from pay-per-crawl (HTTP 402), still in private beta.

Last piece: the Content Signals Policy (Sep 24, 2025) adds signals to the managed robots.txt — Content-Signal: search=yes, ai-train=no — already on more than 3.8 million domains. But these are signals, not enforcement. The wall is Robotcop (December 2024): network-level robots.txt enforcement against the bot that declares one thing and does another. A signal for the one that complies; a wall for the one that doesn't.

Crawler Hints: Cloudflare already speaks IndexNow

Cloudflare is already an IndexNow submitter. Crawler Hints (Cache → Configuration, free on all plans) notifies search engines when it detects changed content, using the IndexNow protocol. It requires proxied records and excludes URLs that respond 4xx or worse.

The nuance: it fires on cache MISS, a passive heuristic: Cloudflare infers the change from cache state, without knowing what changed or whether it matters. A per-event push is a different thing: the "product updated" webhook carries the exact URL at the exact moment, including removals, which a MISS never represents well. They don't compete: turn on Crawler Hints and add per-event push on top. For stores, that's IndexNow Connect: it listens to your catalog's webhooks (additions, changes, removals) and sends each URL to IndexNow instantly, with no heuristics and without touching your store's code.

Workers: SEO at the edge

A Worker runs before your origin. Two concrete SEO uses:

The free plan includes 100,000 daily requests: for these uses, it's more than enough.

Observability: watch what Cloudflare decides for you

If you do the SEO but don't manage the client's Cloudflare, don't ask for admin: ask for read-only access to Security → Events, Security → Bots and AI Crawl Control. Enough to audit without touching anything.

Final checklist

  1. Apex and www proxied; no A records pointing at a SaaS platform's IPs.
  2. TLS on Full (strict) with an origin certificate.
  3. Aggressive static caching; HTML purge on publish for prices, stock or content.
  4. Skip rule with cf.client.bot first in the order; Security → Events with no blocks on Googlebot/Bingbot.
  5. AI Crawl Control reviewed — mandatory if the zone was born after Jul 1, 2025 — with the search bots on Allow.
  6. robots.txt and Content Signals consistent with what the WAF actually does.
  7. Crawler Hints enabled, plus per-event push for the catalog.
  8. Read-only access agreed for whoever audits.

Frequently asked questions

Does Bot Fight Mode block Googlebot?

It shouldn't: the docs say verified bots are exempt, but there are false positives reported in the community. On the Free plan there are no exceptions or Skip rules, so if you see 403s to Googlebot in Security → Events, your options are turning off Bot Fight Mode or moving to Pro and using Super Bot Fight Mode with a Skip rule.

Does caching HTML improve or hurt SEO?

Both, and your invalidation decides which. It improves the response time users and crawlers see, but stale HTML served from the edge pushes outdated prices and stock into the index and into AI answers. If you can't purge on publish, don't cache HTML.

Should I block GPTBot or not?

It's a business decision about training, not about visibility. Blocking GPTBot doesn't remove you from ChatGPT search — that's governed by OAI-SearchBot, which you should allow. The serious mistake is blocking both by reflex.

Why does Search Console report a 403 if my server responds 200?

Because the 403 isn't emitted by your server: Cloudflare emits it before the request reaches the origin. A curl from your network won't reproduce it because you're not Googlebot. Filter Security → Events by user-agent, identify the rule and create a Skip rule first in the order.

Connect your store for free → — Crawler Hints notifies when the cache guesses; IndexNow Connect notifies when your catalog actually changes.