Is your site blocking AI search without knowing it? Checking your AI crawler settings in autumn 2026
· Technical Guide

Three separate switches now decide whether AI search engines can read your content: your robots.txt, your CDN and Google's own settings. Most sites have deliberately set one of them. Some have set none.
For most of the web's history, "can search engines see my site?" had a one-line answer: check robots.txt. That stopped being true this summer. Between a change to Cloudflare's defaults, a new opt-out from Google and a steady rise in AI crawler traffic, whether ChatGPT, Claude, Perplexity or Google's AI Overviews can reach your pages now depends on several settings that live in different places, owned by different people and often left at their defaults.
This is a practical guide to finding those settings, understanding what each controls and deciding, deliberately, what you want AI systems to be able to do with your content.
What changed this summer
Cloudflare split AI traffic into three categories. Cloudflare now sorts AI bot traffic into Search, Agent and Training. Under its new defaults, which took effect on 15 September, Agent and Training bots are blocked by default on ad-bearing pages for new domains. Search crawlers are treated separately, so a site can allow AI search while refusing model training. Cloudflare sits in front of a very large share of the web, so this one change shifts many sites' behaviour without anyone touching a config file. (Cloudflare blog, Cloudflare changelog)
Google is adding an AI-only opt-out. Google is adding a setting that lets site owners keep their content out of AI Overviews and AI Mode specifically, without blocking regular Search or its AI training crawlers. At the same time, Search Console's generative-AI performance reports, covering AI Overviews, AI Mode and AI in Discover, completed their worldwide rollout on 31 August. (Search Engine Roundtable)
AI crawlers are busier, and most sites haven't decided anything. GPTBot traffic is reportedly up 305% year on year (Passionfruit), yet BrightEdge's tracking finds that only 19% of sites have specific crawl rules for ChatGPT's bots (BrightEdge). The other 81% are running on whatever their CMS, host or CDN happens to do.
The three layers that decide what AI can see
It helps to think of AI access as three layers, checked in this order.
Layer 1: robots.txt (what you ask for)
robots.txt is a request, not a lock, but the major AI companies say they respect it, and each now runs several crawlers with different jobs. The distinction that matters most is search versus training:
- OpenAI:
OAI-SearchBotfetches pages for ChatGPT search results;GPTBotcollects content for model training;ChatGPT-Userfetches a page when a user asks ChatGPT to visit it. (OpenAI crawler documentation) - Google:
Googlebotfeeds Search, including AI Overviews and AI Mode.Google-Extendedis a separate token that controls use of your content for Gemini model training and grounding; it does not remove you from Search or AI Overviews. (Google crawler documentation) - Anthropic and Perplexity follow a similar pattern, with separate agents for training, search indexing and user-initiated fetches. Check each company's current documentation, because the agent names change more often than you'd hope.
The most common mistake we see is a blanket block, often copied from a 2023 "block the AI scrapers" guide, that shuts out the search crawlers along with the training ones. The result is a site that has opted out of being cited while believing it has only opted out of being trained on.
Layer 2: your CDN or firewall (what actually gets through)
This is the layer most site owners forget, because it doesn't live in the codebase. Your robots.txt can welcome OAI-SearchBot while Cloudflare's bot management, a WAF rule or a "block AI bots" toggle turned on during a scraping panic quietly returns a 403.
If you're on Cloudflare, open Security → Bots (the exact labels vary by plan) and check how the Search, Agent and Training categories are set for your zone. If your domain was added recently, the new defaults may apply. On other CDNs and hosts, look for bot-management or "AI crawler" settings, and check your server logs for 403 or 429 responses served to known AI user agents.
Layer 3: platform controls (what gets shown)
Even when a crawler can fetch your page, the platform decides how your content may appear. On Google, that means the existing snippet controls (nosnippet, data-nosnippet and max-snippet), which also govern what can be quoted in AI Overviews, and soon the dedicated AI Overviews/AI Mode opt-out.
The trade-off nobody mentions: opting out has a price
Blocking AI features isn't free. Tracking data published in August shows that Top Stories carousels now render inside AI Overviews for trending US news queries, which means a publisher that opts out of AI features may also lose that placement (Search Engine Land). The line between "AI answer" and "search result" is getting blurrier, and an opt-out designed to protect content may also cut off traffic you actually want.
That doesn't mean everyone should allow everything. Publishers negotiating licensing deals, sites with paywalled content and businesses with genuine scraping problems all have good reasons to be selective. The point is to make the decision on purpose, with the costs in view.
A 20-minute AI access audit
- Read your robots.txt as a crawler would. Visit
yourdomain.com/robots.txtand list every AI-related user agent it mentions. For each one, write down whether it's a search, training or user-fetch crawler, and whether that matches what you intend. - Check the CDN layer. Log in to Cloudflare (or your CDN/host) and note the current AI bot settings. Pay special attention if the domain was set up recently.
- Test with a real user agent. From a terminal, request a key page while identifying as an AI crawler, for example
curl -I -A "OAI-SearchBot" https://yourdomain.com/your-page. A200is good; a403means something between you and the crawler is saying no. (Some firewalls also verify crawler IP ranges, so treat a200here as necessary rather than sufficient.) - Search your server logs for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Googlebot. If a crawler you've allowed never appears, or appears only with error responses, you have a problem.
- Look at Search Console's generative-AI reports. Now that they're available worldwide, they're the most direct evidence of whether your pages are appearing in AI Overviews and AI Mode at all.
- Write down your policy in one sentence, such as "We allow AI search and user fetches, and block training." Then check that all three layers actually implement it.
Once access is sorted, the next question is whether crawlers can read what they fetch. Many can't run JavaScript, which we cover in GPTBot can't see your JavaScript. For the broader checklist, see 7 free ways to check whether AI search engines can actually see your website, and for the case for signposting your content to AI crawlers, the importance of having an llms.txt file.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search? Not on its own. OpenAI uses GPTBot for training and OAI-SearchBot for search, so you can block one and allow the other. But a CDN or firewall rule that blocks AI bots in general may block both.
Does blocking Google-Extended remove me from AI Overviews? No. AI Overviews and AI Mode draw on Google's normal Search index, which is crawled by Googlebot. Google-Extended controls use of your content for Gemini training and grounding. Google is adding a separate control for AI Overviews and AI Mode specifically.
What did Cloudflare change on 15 September 2026? Cloudflare now classifies AI bot traffic as Search, Agent or Training, and blocks Agent and Training bots by default on ad-bearing pages for new domains. Existing settings should be reviewed rather than assumed.
Should I block AI crawlers? It depends on your business model. Blocking training crawlers while allowing search crawlers is a common middle ground. Blocking AI search features entirely can cost visibility, including, for news publishers, Top Stories placements that now appear inside AI Overviews.
The kicker
For twenty-five years, the web's welcome mat was a text file. It still matters, but it now shares the doorway with a CDN dashboard and a Google toggle, and they don't always agree. The sites that come out of this well won't necessarily be the ones that let every bot in. They'll be the ones that know exactly which bots they've let in, and why.
Want to know how your site looks to AI crawlers right now? Run a free CiteSite audit: it checks crawler access, server-side renderability and structured data in one pass.
Sources: Cloudflare: new AI traffic options for all customers; Cloudflare changelog: new options to manage AI traffic; Search Engine Roundtable: September 2026 Google webmaster report; BrightEdge weekly AI search insights; Search Engine Land: AI Overview data from 51,000 tracked events; Passionfruit: JavaScript rendering and AI crawlers; OpenAI crawler documentation; Google common crawlers documentation.