Cloudflare's 'Disallow AI Training' Switch Is Live: Check Your Crawler Settings
💡 Tool Tip:robots.txt Generator, Sitemap Generator, Google Index Checker
On September 15, 2026, Cloudflare changed its crawler controls. The important part is not that a new switch appeared - it is that two old switches changed meaning. Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot and Googlebot, so selecting them affects search as well as training. The tool for refusing training while staying indexed is a new setting called Disallow AI Training. If you run a site behind Cloudflare, this is worth a twelve-minute review.
The default changed - and so did the meaning of the switch you ticked months ago
1. What Changed on September 15
Cloudflare's own post lists four changes. First, Block and Block on pages with ads now apply to all training crawlers including mixed-use crawlers - Applebot, Bingbot and Googlebot among them - so either setting impacts search as well as training; to stop training and keep search, use Disallow AI Training. Second, the 'Block AI Bots' toggle is deprecated in favour of the more granular Search, Training and Agent controls. Third, Managed Robots.txt is deprecated in favour of Bot Preference Sync, with existing users migrated. Fourth, Disallow AI Training becomes part of the recommended configuration for certain new domains. The migration note is the actionable part: in almost every case customers had nothing to do, because their preferences carried over automatically.
# 1. What changed on 2026-09-15, in Cloudflare's own words
CHANGES = [
"Block and Block on pages with ads now apply to mixed-use crawlers, "
"including Applebot, Bingbot and Googlebot, so either setting impacts "
"search as well as training.",
"Block AI Bots is deprecated in favour of the more granular Search, "
"Training and Agent controls.",
"Managed Robots.txt is deprecated in favour of Bot Preference Sync; "
"customers who enabled it migrate to the new system.",
"Disallow AI Training becomes part of the recommended configuration for "
"certain new domains.",
]
SOURCE = "blog.cloudflare.com/accountable-mixed-use-ai-crawlers (2026-09-15)"
# The migration note is the actionable bit: existing customers' preferences
# were carried over automatically. Nothing to do - unless you were relying on
# Block to mean 'block the training crawler only'.2. Three Behaviours, Four Settings, One That Keeps Search
Cloudflare classifies bots by behaviour. Search is crawling to build a search index. Training is crawling to train or fine-tune a model. Agent is a user-directed agent visiting a page on behalf of a human, such as chat fetch bots and browser-use agents. On top of that sit four settings. Allow lets everything through unless another setting or a WAF rule blocks it. Disallow AI Training publishes the applicable no-training preference in robots.txt; Accountable mixed-use crawlers remain allowed for search while every other training crawler is blocked. Block on pages with ads blocks crawlers - including mixed-use crawlers - only on pages detected to be serving an ad. Block blocks everything, mixed-use crawlers included. Two details matter: Disallow AI Training is only available as a setting for Training, not Search or Agent; and there is no 'Disallow AI Training on pages with ads', because Cloudflare can detect ad pages but that list is too large and changes too often to enumerate in robots.txt.
# 2. Three behaviours, four settings - and only one keeps search
BEHAVIOURS = {
"Search": "crawling to build a search index",
"Training": "crawling to train or fine-tune a model",
"Agent": "a user-directed agent visiting a page on behalf of a human",
}
SETTINGS = {
"Allow": "all crawlers allowed unless another setting or WAF rule blocks them",
"Disallow AI Training": "publishes the no-training preference in robots.txt; "
"Accountable mixed-use crawlers stay allowed for search; "
"every other training crawler is blocked",
"Block on pages with ads": "crawlers, including mixed-use crawlers, "
"blocked only on pages detected to serve an ad",
"Block": "all crawlers, including mixed-use crawlers, are blocked",
}
# Disallow AI Training is a Training-only setting. There is no equivalent for
# Agent, because standards such as ai-prefs have not matured yet, and no
# 'Disallow on pages with ads', because that list cannot be expressed in robots.txt.Granular control starts with separating who is crawling from why
3. The Two Numbers Behind the Redesign
Two figures in the post explain the whole design. Fewer than 1% of Cloudflare sites choose to block Search bots. Meanwhile 17% of sites choose to enable some mechanism to block training. Almost every site owner treats search as beneficial; training is a different question entirely. The old one-size-fits-all 'Block AI' approach forced owners to give up discoverability in order to refuse training. Cloudflare also argues a robots.txt directive alone cannot solve this: anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it. A network, in its account, can do all four - publish the preference, identify the crawler, classify the intent, block those that ignore it, and report what each operator actually does.
// 3. The numbers Cloudflare published to justify granular controls
const adoption = {
sitesBlockingSearch: "<1%",
sitesBlockingTrainingSomehow: "17%",
interpretation: [
"almost every site owner treats search as beneficial",
"training is a different question entirely",
"a one-size-fits-all 'Block AI' switch forced owners to give up one to refuse the other",
],
};
const whyRobotsTxtAloneIsNotEnough = [
"anyone can publish a directive",
"it cannot identify who is crawling",
"it cannot determine why they are crawling",
"it cannot stop a crawler that ignores it",
];
// Cloudflare's claim is that a network can do all four: publish the preference,
// identify the crawler, classify the intent, block those that ignore it, and
// report what each operator actually does.4. What 'Accountable' Means
To distinguish mixed-use crawler operators that respect site owner choice, Cloudflare created the Accountable designation. To qualify, an operator must meet or commit to meeting four requirements: a mechanism to opt out of AI training via robots.txt or a similar standard; a mechanism to opt out of AI summaries directly with the operator, and through Cloudflare next year; URL-level visibility into which pages were made available for training plus metrics on how content appeared in search; and assurance that opting out of training will not affect traditional search results. Apple, Google and Microsoft currently qualify. Per crawler: Applebot opts out of training via a robots.txt Disallow for Applebot-Extended, Googlebot via Google-Extended, and Bingbot currently expresses training preferences through the NOARCHIVE meta tag with robots.txt-level support targeted for early 2027. Cloudflare also categorises crawlers from Amazon, Anthropic, Meta and OpenAI as Accountable, noting those organisations separate search and training crawlers so training can be blocked without affecting search.
# 4. What 'Accountable' means for each mixed-use crawler
ACCOUNTABLE_REQUIREMENTS = [
"a mechanism to opt out of AI training via robots.txt or a similar standard",
"a mechanism to opt out of AI summaries with the operator directly, "
"and through Cloudflare next year",
"URL-level visibility into which pages were made available for training, "
"plus metrics showing how content appeared in search",
"assurance that opting out of AI training will not affect traditional search results",
]
PER_CRAWLER = {
"Applebot": "opt out of training via a robots.txt Disallow for Applebot-Extended",
"Googlebot": "opt out of training via a robots.txt Disallow for Google-Extended",
"Bingbot": "training preferences currently expressed through the NOARCHIVE meta tag; "
"robots.txt-level support targeted for early 2027",
}
# Cloudflare also categorises the relevant crawlers from Amazon, Anthropic, Meta
# and OpenAI as Accountable, noting they separate search and training crawlers.The robots.txt actually served is the only version that counts
5. The Four Steps to Check Today
First, inventory: list the crawlers actually hitting your origin, by user agent and by volume. Second, decide, per behaviour: keep Search allowed, since it is the channel that brings humans; use Disallow AI Training for Training unless you license your content; and treat Agent separately, because agents fetch a page with nobody there to see the ads. Third, verify: fetch your live robots.txt and read the directives that are actually served, confirm whether Cloudflare is injecting anything at the edge, and re-check after any plan or zone setting change. Fourth, monitor referral traffic from AI assistants, not just classic organic. Three dates are worth putting in the calendar: when the AI summaries opt-out launches, when robots.txt-level training preferences ship for Bing, and whenever your monetisation model changes.
{
"crawler_policy_review": {
"date": "2026-09-20",
"step_1_inventory": "list the crawlers actually hitting your origin, by user agent and by volume",
"step_2_decide": {
"search": "keep allowed - it is the channel that brings humans",
"training": "usually Disallow AI Training, unless you license your content",
"agents": "decide separately; agents fetch the page with nobody there to see the ads"
},
"step_3_verify": [
"fetch your live robots.txt and read the directives that are actually served",
"confirm whether Cloudflare is injecting directives at the edge",
"re-check after any plan or zone setting change"
],
"step_4_monitor": "watch referral traffic from AI assistants, not just classic organic",
"revisit_when": [
"the AI summaries opt-out launches",
"robots.txt-level training preferences ship for Bing",
"your monetisation model changes"
]
}
}📌 Frequently Asked Questions
What exactly changed on September 15, 2026?
Per Cloudflare's post: Block and Block on pages with ads now apply to mixed-use crawlers including Applebot, Bingbot and Googlebot, so they affect search as well as training; Block AI Bots is deprecated; Managed Robots.txt is replaced by Bot Preference Sync; and Disallow AI Training becomes a recommended setting for certain new domains. Existing customers' settings migrated automatically.
I want to block training but stay in search. Which setting should I use?
Use Disallow AI Training. It publishes the no-training preference in robots.txt, keeps Accountable mixed-use crawlers allowed for search, and blocks every other training crawler.
Why is there no 'Disallow AI Training on pages with ads'?
Cloudflare explains that Disallow AI Training works by publishing a preference in robots.txt, while the set of ad-serving pages is too large and changes too frequently to enumerate there. Block on pages with ads is the setting for that dimension.
What does the Accountable designation mean?
It recognises mixed-use crawler operators that meet or commit to four requirements: an opt-out from AI training via robots.txt or a similar standard, an opt-out from AI summaries, URL-level visibility plus search metrics, and assurance that opting out of training will not affect search ranking. Apple, Google and Microsoft currently qualify.
What happens to sites that had the old 'Block AI Bots' setting enabled?
Cloudflare's migration table maps the legacy setting onto the new controls: a previous Block becomes Search set to Allow, Training set to Disallow AI Training, and Agent set to Block on pages with ads. In almost every case no action is required, but it is worth confirming what is actually live.
🔧 Recommended Tools
robots.txt Generator
Write your crawler policy as a robots.txt you can actually verify
Sitemap Generator
Keep discovery healthy so a bot policy does not starve search
Google Index Checker
Confirm pages are still indexed after you change settings
User-Agent Parser
Identify who is actually hitting your origin
Meta Tag Analyzer
Check that nosnippet and NOARCHIVE controls behave as intended
📚 Sources
- Cloudflare Blog (2026-09-15) - Have it both ways: stay discoverable in search while disallowing AI training: the Disallow AI Training setting, the September 15 changes, the Accountable designation and its four requirements, per-crawler opt-outs for Applebot, Googlebot and Bingbot, and the migration tables
- Cloudflare Changelog (2026-07-01) - New options to manage AI traffic: the Search / Agent / Training behaviours, the three block presets, and the September 15 default for new domains
- Cloudflare Blog (2026-07-01) - Your site, your rules: new AI traffic options for all customers, including Free