Cloudflare 的「Disallow AI Training」已生效:请检查你的爬虫设置
💡 工具推荐:robots.txt 生成器, 站点地图生成器, Google 收录检查
2026 年 9 月 15 日,Cloudflare 对自己的爬虫控制做了一次语义层面的改动。这一天的影响不在于新增了开关,而在于旧开关的含义被改写了:Block 与 Block on pages with ads 从此也作用于「混合用途爬虫」,其中包括 Applebot、Bingbot 与 Googlebot——也就是说,勾选它们会在阻止训练的同时影响搜索收录。真正用于「拒绝训练但保持收录」的工具,是一个新的设置:Disallow AI Training。如果你有站点在 Cloudflare 上,这篇文章值得花十二分钟核一遍。
默认设置变了,你上次点过的那个开关含义也变了
一、9 月 15 日到底改了什么
按 Cloudflare 官方博客,9 月 15 日的变更共四点。第一,Block 与 Block on pages with ads 现在适用于混合用途爬虫,包括 Applebot、Bingbot 与 Googlebot,因此这两个设置会同时影响搜索与训练;想在阻止训练的同时保住搜索,应当使用 Disallow AI Training。第二,「Block AI Bots」被弃用,让位于更细粒度的 Search、Training、Agent 三类控制。第三,「Managed Robots.txt」被弃用,由 Bot Preference Sync 取代,启用过的客户会迁移到新系统。第四,Disallow AI Training 将成为特定新域名的推荐配置。公告同时说明:几乎所有情况下客户无需动作,原有设置会自动迁移。
# 1. What changed on 2026-09-15, in Cloudflare's own words
CHANGES = [
"Block and Block on pages with ads now apply to mixed-use crawlers, "
"including Applebot, Bingbot and Googlebot, so either setting impacts "
"search as well as training.",
"Block AI Bots is deprecated in favour of the more granular Search, "
"Training and Agent controls.",
"Managed Robots.txt is deprecated in favour of Bot Preference Sync; "
"customers who enabled it migrate to the new system.",
"Disallow AI Training becomes part of the recommended configuration for "
"certain new domains.",
]
SOURCE = "blog.cloudflare.com/accountable-mixed-use-ai-crawlers (2026-09-15)"
# The migration note is the actionable bit: existing customers' preferences
# were carried over automatically. Nothing to do - unless you were relying on
# Block to mean 'block the training crawler only'.二、三种行为、四个设置,只有一个能保住搜索
Cloudflare 按行为分类爬虫:Search 是「为建立搜索索引而抓取」;Training 是「为训练或微调模型而抓取」;Agent 是「代表人类实时访问页面的用户指令型智能体」。在此之上有四个设置:Allow(除非其他设置或 WAF 规则拦截,否则全部放行);Disallow AI Training(把「不训练」的偏好写进 robots.txt,被评为 Accountable 的混合用途爬虫仍可抓取以服务搜索,其余所有训练类爬虫一律拦截);Block on pages with ads(仅在检测到展示广告的页面上拦截,含混合用途爬虫);Block(拦截全部爬虫,含混合用途爬虫)。两个细节值得记住:Disallow AI Training 只能作为 Training 的设置,Search 与 Agent 没有对应选项;也不存在「仅在广告页 Disallow AI Training」,因为那个页面清单太大、变化太快,无法用 robots.txt 枚举。
# 2. Three behaviours, four settings - and only one keeps search
BEHAVIOURS = {
"Search": "crawling to build a search index",
"Training": "crawling to train or fine-tune a model",
"Agent": "a user-directed agent visiting a page on behalf of a human",
}
SETTINGS = {
"Allow": "all crawlers allowed unless another setting or WAF rule blocks them",
"Disallow AI Training": "publishes the no-training preference in robots.txt; "
"Accountable mixed-use crawlers stay allowed for search; "
"every other training crawler is blocked",
"Block on pages with ads": "crawlers, including mixed-use crawlers, "
"blocked only on pages detected to serve an ad",
"Block": "all crawlers, including mixed-use crawlers, are blocked",
}
# Disallow AI Training is a Training-only setting. There is no equivalent for
# Agent, because standards such as ai-prefs have not matured yet, and no
# 'Disallow on pages with ads', because that list cannot be expressed in robots.txt.把「谁在爬、为什么爬」分开,才谈得上精细控制
三、Cloudflare 给出的两个数字,解释了为什么要做细粒度控制
公告里有两条数据很有说服力:在 Cloudflare 的站点中,选择阻止 Search 爬虫的不到 1%;而选择启用某种机制阻止训练的比例达到 17%。换句话说,几乎所有站长都认为搜索是有益的,但训练完全是另一个问题。而在过去,一个「Block AI」的一刀切开关会逼站长二选一:要么允许训练,要么牺牲可发现性。公告还解释为什么单靠 robots.txt 不够:任何人都能发布指令,但它无法识别「谁在爬」、无法判断「为什么爬」、也无法阻止无视指令的爬虫——而网络层可以做到这四件事:发布偏好、识别主体、判定意图、拦截无视者,并通过 Radar 报告各运营方的实际行为。
// 3. The numbers Cloudflare published to justify granular controls
const adoption = {
sitesBlockingSearch: "<1%",
sitesBlockingTrainingSomehow: "17%",
interpretation: [
"almost every site owner treats search as beneficial",
"training is a different question entirely",
"a one-size-fits-all 'Block AI' switch forced owners to give up one to refuse the other",
],
};
const whyRobotsTxtAloneIsNotEnough = [
"anyone can publish a directive",
"it cannot identify who is crawling",
"it cannot determine why they are crawling",
"it cannot stop a crawler that ignores it",
];
// Cloudflare's claim is that a network can do all four: publish the preference,
// identify the crawler, classify the intent, block those that ignore it, and
// report what each operator actually does.四、「Accountable」这个标签意味着什么
为了区分「愿意尊重站长选择」的混合用途爬虫,Cloudflare 设立了 Accountable 认定,要求运营方满足或承诺满足四项:提供通过 robots.txt 或类似标准退出 AI 训练的机制;提供直接向运营方退出 AI 摘要的机制(明年将通过 Cloudflare 提供);提供 URL 级别的可见性——哪些页面被用于训练,以及内容在搜索中如何呈现的指标;保证退出 AI 训练不会影响传统搜索结果。目前 Apple、Google、Microsoft 都符合资格。具体到爬虫:Applebot 通过 robots.txt 中对 Applebot-Extended 的 Disallow 退出训练;Googlebot 通过 Google-Extended 的 Disallow;Bingbot 目前通过 NOARCHIVE meta 标签表达训练偏好,域名/站点级的 robots.txt 支持目标定在 2027 年初。Cloudflare 同时把 Amazon、Anthropic、Meta、OpenAI 的相关爬虫归为 Accountable,理由是这些组织分离了搜索与训练爬虫,因此可以只拦截训练爬虫而不影响搜索。
# 4. What 'Accountable' means for each mixed-use crawler
ACCOUNTABLE_REQUIREMENTS = [
"a mechanism to opt out of AI training via robots.txt or a similar standard",
"a mechanism to opt out of AI summaries with the operator directly, "
"and through Cloudflare next year",
"URL-level visibility into which pages were made available for training, "
"plus metrics showing how content appeared in search",
"assurance that opting out of AI training will not affect traditional search results",
]
PER_CRAWLER = {
"Applebot": "opt out of training via a robots.txt Disallow for Applebot-Extended",
"Googlebot": "opt out of training via a robots.txt Disallow for Google-Extended",
"Bingbot": "training preferences currently expressed through the NOARCHIVE meta tag; "
"robots.txt-level support targeted for early 2027",
}
# Cloudflare also categorises the relevant crawlers from Amazon, Anthropic, Meta
# and OpenAI as Accountable, noting they separate search and training crawlers.线上生效的 robots.txt,才是唯一算数的版本
五、今天该核对的四步
第一,盘点:按 user agent 与流量列出真正在打你源站的爬虫。第二,决策:Search 通常保持允许(它是带人来的渠道);Training 通常是 Disallow AI Training,除非你本来就授权内容被训练;Agent 需要单独决策,因为它抓取页面时并没有人在看广告。第三,验证:抓取你线上的 robots.txt,读实际返回的指令,确认 Cloudflare 是否在边缘注入了额外指令,并且在任何套餐或 zone 设置变更后重新检查。第四,监控:除了传统自然流量,也要盯来自 AI 助手的引荐流量。另外记三个复查触发点:AI 摘要的退出机制上线时、Bing 的 robots.txt 级训练偏好在 2027 年初发布时、以及你自身变现模式变化时。
{
"crawler_policy_review": {
"date": "2026-09-20",
"step_1_inventory": "list the crawlers actually hitting your origin, by user agent and by volume",
"step_2_decide": {
"search": "keep allowed - it is the channel that brings humans",
"training": "usually Disallow AI Training, unless you license your content",
"agents": "decide separately; agents fetch the page with nobody there to see the ads"
},
"step_3_verify": [
"fetch your live robots.txt and read the directives that are actually served",
"confirm whether Cloudflare is injecting directives at the edge",
"re-check after any plan or zone setting change"
],
"step_4_monitor": "watch referral traffic from AI assistants, not just classic organic",
"revisit_when": [
"the AI summaries opt-out launches",
"robots.txt-level training preferences ship for Bing",
"your monetisation model changes"
]
}
}📌 常见问题 FAQ
2026 年 9 月 15 日 Cloudflare 究竟改了什么?
据官方博客:Block 与 Block on pages with ads 开始适用于混合用途爬虫(含 Applebot、Bingbot、Googlebot),会同时影响搜索;Block AI Bots 被弃用;Managed Robots.txt 被 Bot Preference Sync 取代;Disallow AI Training 成为特定新域名的推荐配置。已有客户设置会自动迁移。
我想阻止训练但保留搜索收录,该用哪个设置?
使用 Disallow AI Training。它会把「不训练」偏好写进 robots.txt;被评为 Accountable 的混合用途爬虫仍可为搜索抓取,其他所有训练类爬虫一律被拦截。
为什么没有「仅在广告页 Disallow AI Training」?
公告解释:Disallow AI Training 通过写入 robots.txt 生效,而「哪些页面展示广告」的清单太大、变化太快,无法用 robots.txt 枚举。广告页维度的控制是用 Block on pages with ads 表达的。
「Accountable」代表什么?
它是 Cloudflare 对爬虫运营方的认定:需满足或承诺满足四项要求——提供退出 AI 训练的机制、提供退出 AI 摘要的机制、提供 URL 级可见性与搜索表现指标、并保证退出训练不影响传统搜索排名。Apple、Google、Microsoft 目前符合资格。
启用过旧的「Block AI Bots」的站点会怎样?
按公告的迁移说明,旧设置会被映射到新控制:曾经的 Block 会迁移为 Search 保持 Allow、Training 变为 Disallow AI Training、Agent 为 Block on pages with ads。几乎所有情况下站长无需手动操作,但建议核对一次实际生效的配置。
🔧 推荐工具
📚 参考资料
- Cloudflare Blog (2026-09-15) - Have it both ways: stay discoverable in search while disallowing AI training: the Disallow AI Training setting, the September 15 changes, the Accountable designation and its four requirements, per-crawler opt-outs for Applebot, Googlebot and Bingbot, and the migration tables
- Cloudflare Changelog (2026-07-01) - New options to manage AI traffic: the Search / Agent / Training behaviours, the three block presets, and the September 15 default for new domains
- Cloudflare Blog (2026-07-01) - Your site, your rules: new AI traffic options for all customers, including Free