Cloudflare's new AI crawler rules: stay found, keep control
Cloudflare now lets you refuse AI training without disappearing from search. Here is what changed, why the old controls cost brands visibility, and why letting crawlers in is still good for AEO.
Published 16 September 2026
At a glance
- Cloudflare's new Disallow AI Training setting blocks training crawlers while Google, Bing and Apple keep indexing you for search.
- The old choice was all or nothing, because the biggest crawlers do search and training at once. Blocking one blocked both.
- AI assistants can only quote pages their crawlers can read, so blocking search or agent bots quietly removes you from answers.
- For most B2B brands the right setting is still to allow everything, then check robots.txt and firewall rules for accidental blocks.

On 15 September 2026, Cloudflare changed how its AI crawler controls work. Until now, a site owner who wanted to stop AI companies training on their content had one blunt tool, and it risked switching off search along with it. The new setting separates the two: you can say no to training and still be read by the search and answer engines your buyers use. For most B2B brands, the right move is simpler still: let the bots in, and check that nothing on your site is quietly keeping them out.
What did Cloudflare actually change?
Cloudflare now sorts crawlers by why they are visiting, not just who sent them. Every bot falls into one of three purposes:
- Search: building an index so pages can be found and quoted in results, including AI answers.
- Training: collecting content to build or fine-tune an AI model.
- Agent: fetching a page because a person asked an assistant to do something right now.
Alongside that, there is a new option called Disallow AI Training. It blocks crawlers that collect content for training, while letting the big search crawlers keep indexing your site. It is available on every Cloudflare plan, in the security settings for each domain.
| Setting | What it does |
|---|---|
| Allow | Every crawler is welcome, unless you block it somewhere else |
| Disallow AI Training | Training crawlers are blocked; accountable search crawlers keep indexing |
| Block on pages with ads | Crawlers are blocked only on pages that carry advertising |
| Block | Everything is blocked, search included |
Why was the old approach a problem?
Because the biggest crawlers do two jobs at once. Googlebot, Bingbot and Applebot each index the web for search and feed their company's AI work. Cloudflare calls these mixed-use crawlers. If you blocked the bot, you lost both. If you allowed it, you accepted both.
That left site owners with a false choice: give your content away for training, or disappear from search. Cloudflare's own numbers show how people resolved it. Fewer than 1% of sites on its network block search bots, because nobody wants to vanish from search. But 17% had switched on some way of blocking AI training.
The blunt tools had side effects that were easy to miss. A "block AI" toggle or a copied-and-pasted robots.txt rule could stop the very assistants you want recommending you, not just the ones building models. Many teams never noticed, because a blocked AI crawler causes no error, no drop in rankings and no alert. You simply stop being mentioned, and nothing tells you why.
What does "accountable" mean here?
Cloudflare will only let a mixed-use crawler through the Disallow AI Training setting if its operator agrees to four things:
- A way to opt out of training, through robots.txt or an equivalent.
- A way to opt out of AI summaries of your pages.
- Visibility, down to individual URLs, of how your content is used in search and training.
- A commitment that opting out of training will not hurt your search rankings.
Apple, Google and Microsoft are named as accountable operators of mixed-use crawlers. Amazon, Anthropic, Meta and OpenAI are named for running separate crawlers for separate jobs, which is the cleaner arrangement Cloudflare would like everyone to reach.
In practice, the opt-outs look like this today:
- Google: add a rule for
Google-Extendedin robots.txt. Google confirms this does not affect search ranking. - Apple: add a rule for
Applebot-Extendedin robots.txt. - Microsoft: use the
NOARCHIVEmeta tag or the Block URLs tool in Bing Webmaster Tools, with robots.txt support planned for early 2027.
Cloudflare also says it will publish which operators are keeping their commitments, through Cloudflare Radar. A promise you can check is worth more than one you cannot.
What happens to my existing settings?
For most sites, nothing you need to act on. Existing customers are being moved across automatically, and the effect is meant to stay the same:
- If you had a training setting of Block or Block on pages with ads, it becomes Disallow AI Training.
- If you blocked everything, including search, that stays blocked.
New domains added from 15 September get sensible defaults. Sites without advertising start with search, training and agents all allowed. Sites that carry ads start with training disallowed and agents blocked on pages with ads.
It is still worth a two-minute look, especially if someone set this up a year ago and has since moved on.
Why is being crawled good for AEO?
Because an assistant can only recommend what it can read. Answer engine optimisation (AEO) is the work of getting your brand named when someone asks ChatGPT, Gemini, Claude or Google's AI Overviews a question in your category. Every one of those answers starts with a crawler visiting a page.
There are three moments where that matters:
- When the answer is written. Most AI answers about buying decisions are grounded in a live search. If the search crawler never read your pricing page, your case studies or your comparison guide, the assistant has nothing of yours to quote. It will quote a competitor instead.
- When someone asks an agent to check. A buyer who says "look at their website and tell me if they do X" sends an agent to your site. Block it, and the assistant reports back that it could not find out.
- When the model learns what exists. Training shapes what a model already knows before it searches: which companies exist in a category and what they are known for. For a B2B brand that wants to be recognised, that baseline awareness is an asset, not a leak.
The size of the prize is real. Cloudflare cites Pew Research finding that more than half of consumers read AI summaries in search, and that they are over 40% more likely to end their search after reading one. In other words, the summary is increasingly the whole visit. And it cites SimilarWeb and QuickSEO data showing that visitors referred by AI search convert at between three and five times the rate of traditional search visitors. Fewer clicks, but much better ones, and only if you were in the answer.
If a crawler cannot read your page, no amount of good writing on that page will get you quoted.
So should I disallow AI training or not?
It depends on how your site earns its keep.
- If you sell a product or service, and your content exists to win customers, allowing everything is usually right. Your articles are marketing. You want as many engines as possible to know them, repeat them and point people to you.
- If your content is the product, as it is for a publisher or a paid research firm, Disallow AI Training is a genuinely good new option. You keep your search traffic and stop your work being absorbed into a model for free.
- Whatever you choose, keep search and agents open. Blocking those is what costs you visibility, and it is the mistake the old controls made easy.
The shift here is from "block AI" to "decide what each visitor is for". That is a healthier way to run a website, and it means you no longer have to trade being found for keeping control.
What should I check this week?
- Open your Cloudflare security settings and confirm what each domain is set to for search, training and agents.
- Read your robots.txt. Look for rules aimed at
GPTBot,ClaudeBot,PerplexityBot,OAI-SearchBot,Google-ExtendedorApplebot-Extended, and make sure each one is there on purpose. - Check your CDN or firewall rules for bot-blocking that sits outside Cloudflare's crawler setting, such as blanket challenges on unknown user agents.
- Ask the engines. Put the questions your buyers ask to ChatGPT, Gemini, Claude and Google, and see whether you are named. A crawler problem shows up here first.
Vizebel checks all of this as part of reading your site: whether each AI crawler can reach your pages, and whether the engines name you when your buyers ask.
Frequently asked questions
Will disallowing AI training hurt my Google rankings?
Not according to Google. Opting out through Google-Extended is confirmed not to affect search ranking or ranking signals, and a commitment like that is one of the conditions Cloudflare requires before treating a crawler as accountable. Apple and Microsoft have given the same assurance for their opt-outs.
Does Disallow AI Training stop my brand appearing in ChatGPT or AI Overviews?
It should not, as long as your search and agent settings stay open. Those answers are largely built from search crawlers and live page fetches, which the new setting is designed to let through. Blocking search or agents is what removes you from answers.
I turned on "block AI bots" last year. Do I need to do anything?
Probably not. Cloudflare is migrating existing training blocks to Disallow AI Training automatically, which keeps search working. It is still worth confirming the result, and checking your robots.txt for older rules that block AI search crawlers directly.
Is letting AI companies train on my content a bad idea?
Not for most businesses that use content to win customers. Being part of what a model already knows makes it more likely to recognise your brand and include you in a shortlist. It is a different calculation if your content is itself what you sell.
When will I be able to control AI summaries separately?
Cloudflare expects to add AI summary controls by early 2027, including limits more nuanced than a simple yes or no. Until then, each operator offers its own route, such as the nosnippet directive for Apple and a setting in Google's webmaster tools.

Written by Olly Barrett
Read next
Why isn't an AI visibility dashboard enough?
Measuring is the cheapest part of the job and the part that changes nothing on its own. What to ask a tool before you buy it.
AEO vs SEO: what carries over and what doesn't
Search optimisation makes a page worth clicking. Answer optimisation makes it worth quoting. Here is which of your existing habits still help, and which now cost you.
How AI decides which brands to name
The five patterns shared by brands that keep getting named in AI answers: being legible, answering the real question, existing in more than one place, and two more.
Ready to find out where you stand?
Give us the address. We’ll read the site, work out your audiences and show you the questions before anything runs. Credits only get spent when you say so.
