On July 1, 2025, Cloudflare changed its default setting for all new domains: block all AI crawlers. Cloudflare protects roughly 20% of all websites on the internet. That means a significant portion of the web shifted overnight from AI can crawl by default to AI needs explicit permission. If your site was added to Cloudflare after that date, or if someone on your team enabled this setting on your existing domain, you may be invisible in ChatGPT, Perplexity and Claude right now. And your server logs will never tell you.
Why This Is Invisible in Your Usual Tools
This is what makes the problem particularly insidious. When Cloudflare blocks an AI crawler, it does so at the edge, before the request ever reaches your server. Your server sees nothing. Your access logs show nothing. GA4 sees nothing. The 403 response is sent silently by Cloudflare, and the only traces appear in Cloudflare's own security dashboard, which most marketing teams never check.
The result: you publish GEO-optimized content, you work on structure, sources, schema markup, and AI engines simply cannot read your pages. The crawler receives a 403 error and moves to the next source. Your competitor, who does not use Cloudflare or has properly configured their access, gets cited in your place.
The Critical Distinction Almost Everyone Misses
There are two fundamentally different types of AI crawlers, and the confusion between them is the root cause of most configuration errors. Training crawlers collect your content to build or refine language models. They have zero impact on what ChatGPT or Perplexity answer your prospects in real time. Retrieval crawlers allow AI engines to fetch sources in real time to cite in their responses. Blocking these is what makes you invisible in AI results.
The most common confusion involves OpenAI. GPTBot is OpenAI's training crawler: you can block it with zero impact on your ChatGPT visibility. OAI-SearchBot is the retrieval crawler used for ChatGPT Search responses: if you block it, your site disappears entirely from ChatGPT answers. Both bots belong to OpenAI but serve opposite functions. Many teams block GPTBot thinking they are protecting their content, without realizing they have also blocked OAI-SearchBot, erasing their ChatGPT visibility in the process.
The Complete List of AI Crawlers to Allow in 2026
To maintain your AI citation visibility while protecting your content from unauthorized training, here are the retrieval crawlers to explicitly allow. For ChatGPT Search: OAI-SearchBot and ChatGPT-User. For Perplexity: PerplexityBot. For Claude: ClaudeBot, Claude-SearchBot and Claude-User. Important: the legacy identifiers Claude-Web and anthropic-ai are deprecated and no longer active in 2026. If your robots.txt still uses those strings, your rules do not apply to Anthropic's current crawlers.
The training crawlers you can block without affecting your AI visibility: GPTBot (OpenAI training), CCBot (Common Crawl training), Google-Extended (Gemini training, independent from Google Search), and Applebot-Extended (Apple AI training). Blocking Google-Extended does not remove you from Google Search or Google AI Overviews: those use standard Googlebot, a separate system.
How to Check Your Cloudflare Configuration in 5 Minutes
Log into your Cloudflare dashboard. Go to Security, then Bots, then AI Crawl Control. If the Block on all pages setting is enabled, all your AI crawlers are blocked. You will also see in the Crawlers tab a list of bots that have attempted to access your site, the number of requests, and whether they were allowed or blocked. This is the source of truth on what AI engines can see from your site.
A second independent check: load your robots.txt directly in a browser by going to yourdomain.com/robots.txt and look for directives on OAI-SearchBot, PerplexityBot and ClaudeBot. If you see Disallow: / for these agents, your AI visibility is cut off. Note: Cloudflare can inject its own directives into your robots.txt at the edge, meaning the file served to crawlers may differ from the one on your server. Always check the publicly served version.
The Ideal robots.txt Configuration for 2026
The goal is to allow retrieval crawlers (which generate your AI citations) while blocking training crawlers (which exploit your content without compensation). Allow OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, Claude-SearchBot, Claude-User, Googlebot and Bingbot. Block GPTBot, CCBot, Google-Extended and Applebot-Extended if you want to protect your content from unauthorized training use.
Important: on Cloudflare, the managed AI Block rule takes precedence over your robots.txt. Even if your robots.txt allows OAI-SearchBot, Cloudflare can block it at the edge before the crawler reads your file. Both configurations must be aligned. Always start by disabling or configuring AI Crawl Control in the Cloudflare dashboard first, then adjust your robots.txt to match.
Verifying That Your Fixes Actually Work
After changing your settings, wait 24 to 48 hours. AI crawlers generally re-read robots.txt before each crawl session, but Cloudflare configuration changes take effect immediately. You can verify by checking Bot Analytics logs in Cloudflare in the days following the fix: you should see the relevant crawlers switch from Blocked to Allowed.
The ultimate and most complete verification: query the LLMs directly on your target prompts and see whether your domain starts appearing in their responses. This is the only method that confirms the crawler authorization is actually translating into citations. Bots can crawl your site without necessarily citing it: configuration is necessary but not sufficient. Content quality remains the final arbiter.
Cloudflare Correctly Configured but Still Invisible in AI: What to Do
If your crawlers are allowed and you are still not cited in ChatGPT or Perplexity on your target queries, the problem is no longer technical. It is a content and citation authority problem. Cloudflare configuration is the necessary condition: without it, nothing works. But it is not sufficient to earn citations. Your content also needs to be structured for extractability, your topical authority needs to be strong enough, and you need to be present on the third-party sources LLMs consult first.
Measure Your Real LLM Visibility with Vizible AI
Fixing your Cloudflare configuration is the first step. But how do you know whether you are actually being cited in ChatGPT, Perplexity or Gemini on the topics where your buyers are researching? The only reliable method is to query the models directly on your target prompts. That is what Vizible AI does: the platform automatically sends your prompts to ChatGPT, Gemini, Perplexity, Claude, Mistral, DeepSeek and Groq every day, records whether your brand is mentioned, at what position, and identifies who is cited in your place when you are not. It is the only way to confirm that your technical configuration is actually translating into visibility in AI responses.
Frequently Asked Questions
How do I know if Cloudflare is blocking my AI crawlers?
Log into your Cloudflare dashboard, go to Security, Bots, AI Crawl Control. Check whether Block on all pages is enabled and review the Crawlers tab to see which bots are being blocked. Also check your public robots.txt at yourdomain.com/robots.txt for Disallow directives on AI retrieval agents.
Does blocking GPTBot prevent me from appearing in ChatGPT?
No. GPTBot is OpenAI's training crawler, not the real-time search crawler. What determines your ChatGPT Search visibility is OAI-SearchBot. You can block GPTBot with zero impact on your ChatGPT citations, as long as OAI-SearchBot is allowed.
Does blocking Google-Extended remove my site from Google AI Overviews?
No. Google AI Overviews use standard Googlebot, not Google-Extended. Google-Extended controls only whether your content can enter Gemini's training data. The two systems are independent and you can block Google-Extended while remaining visible in AI Overviews.
Can Cloudflare override my robots.txt?
Yes. Cloudflare operates at the edge and can block crawlers before they read your robots.txt. The Cloudflare managed AI Block rule takes precedence over all robots.txt and WAF rules. You must configure both in alignment: disable or configure AI Crawl Control in Cloudflare first, then adjust your robots.txt to match.
My site was created before July 2025. Am I affected?
The Block on all pages default applies to new domains added after July 1, 2025. Existing domains were not automatically changed. However, if someone on your team manually enabled AI Crawl Control, or if you have robots.txt rules blocking AI bots, you may still be invisible. The check is still necessary.
Check Whether AI Engines Can See and Cite You
Fixing Cloudflare takes 10 minutes. Knowing whether it actually worked takes Vizible AI. The platform monitors every day what ChatGPT, Gemini, Perplexity and Claude say about your brand, whether your domain is being cited, and what your competitors are doing that you are not yet doing. Start your free 14-day trial at Vizible AI, no credit card required.




