26% 的研发由 AI 主导:Anthropic 用三项指标打开实验室的门
💡 工具推荐:AI Token 计数器, 文本统计, Markdown 编辑器
2026 年 9 月 17 日,Anthropic 发布了一份不太像公司公告的公告:《理解前沿实验室内部 AI 研发节奏的测量方法》。它没有发布新模型,而是提出三项可被外部跟踪的指标——AI 有多少研发是 AI 自己做的、智能体的行为被监督到什么程度、算力如何分配——并附上了 2026 年 8 月自家的一份快照。这件事的价值不在于数字有多大,而在于它第一次把「实验室内部到底发生了什么」拆成了可以被第三方核查的问题清单。
把实验室内部拆成可核查的问题清单
一、一个能读懂百分比的前置刻度:Epoch AI 的自动化等级
Anthropic 用了 Epoch AI 提出的「自动化等级」(Automation Level,AL)刻度,从 AL0(无 AI 参与)到 AL5(AI 完全自主运行,无人在环)。中间两级是关键:AL3 是「AI 协作」——在人类紧密指导下完成大量工作;AL4 是「AI 主导」——给定高层提示即可端到端完成大部分任务,人类只做监督。据 2026 年 8 月的快照:Claude 在任何被测的 AI 研发子集上都还没有达到完全自主;它「主导」了 Anthropic 全部 AI 研发工作的 26%;达到或超过「AI 协作」这一级的比例高于 90%。把这三个数字放在一起,才能真正理解「26%」意味着什么:不是 AI 在做四分之一的活,而是四分之一的活已经不需要人类给出每一步的方向。
# 1. The scale that makes the number readable: Epoch AI's Automation Level
AUTOMATION_LEVELS = {
"AL0": "no AI involvement",
"AL3": "AI collaborates: large chunks of work under close human direction",
"AL4": "AI leads: most of the task end-to-end from a high-level prompt, human supervises",
"AL5": "AI operates fully autonomously, no human in the loop",
}
# Anthropic R&D Automation Index methodology, in one sentence:
# catalogue every kind of AI R&D work, rate how automated each task is, aggregate.
SNAPSHOT_AUG_2026 = {
"fully_autonomous_subsets": 0, # Claude is at AL5 for no measured subset
"claude_leads": 0.26, # AL4
"at_or_above_collaborates": ">0.90", # AL3+
}
def progress_is_news(prev, cur):
return "automation moved a level, not a percentage point" if prev != cur else "flat"二、为什么单一百分比不能当指标:Anthropic 承认的两个障碍
报告罕见地写清了自己的局限。第一,缺少通用方法论,因此目前各实验室的数字无法直接横向比较。第二,也是最微妙的一点:他们用自己的模型去评估自己的系统,这意味着「裁判」模型可能与被评估的模型共享同样的失败模式。Anthropic 给出的缓解路径有三条:由独立第三方核查;由其他开发者的模型核查,并设置护栏以限制竞争敏感数据的交换;以及把这些测量值当作触发条件,例如在模型被用于进一步 AI 研发之前,强制设置一段固定的测试窗口。对任何要做内部 AI 治理的团队来说,这里可复制的原则很朴素:如果数字无法被验证,那就把方法公开。
// 2. Why a single percentage is not a metric - the two stated obstacles
const reportingGaps = {
methodology: "no common methodology across labs, so numbers are not comparable yet",
judgeBias: "labs evaluate their own systems with their own models, so the judge "
+ "can share the same failure modes as the model under review",
};
const mitigationsAnthropicProposes = [
"verification by an independent third party",
"verification by other developers' models, with guardrails limiting the exchange of competitively sensitive data",
"using the measures as a trigger for stronger requirements, such as a fixed testing window before a new model is used for further AI R&D",
];
// Design takeaway: if you cannot verify the number, publish the method.无法验证的数字,就要连同方法一起公开
三、智能体监督是三个数字,而不是一个
第二项测量针对的是智能体行为的监督,Anthropic 把它拆成三个不同含义的指标:覆盖率(coverage),即一个智能体的动作在执行前后经过监控的比例;复核延迟(review latency),即动作发生到被检查之间的时间,先由自动监控、再由人复核;以及升级率(escalation rate),即被升级处理的智能体活动占比。这个拆法值得直接搬进工程实践:覆盖率回答「有没有被看见」,延迟回答「来得不来得及」,升级率回答「看见了之后有没有人管」。Anthropic 自己给出的动机也很直接——就目前的监控数据看,单个智能体很少越界;但当经济体中运行着数以百万、十亿计的智能体时,罕见事件就会规律地发生。
# 3. Oversight of agents is three different numbers, not one
def oversight_metrics(agent_actions):
return {
# share of an agent's actions that pass through a monitor, before or after execution
"coverage": len([a for a in agent_actions if a.monitored]) / max(len(agent_actions), 1),
# time between an action and its review: first automated, then human
"review_latency": {
"automated_monitor_seconds": percentile(agent_actions, "auto_review_latency", 50),
"human_triage_seconds": percentile(agent_actions, "human_review_latency", 50),
},
# share of agent activity that escalates
"escalation_rate": len([a for a in agent_actions if a.escalated]) / max(len(agent_actions), 1),
}
# Anthropic's framing is the part worth stealing: individual agents rarely
# misbehave, but at millions or billions of agents, rare events happen regularly.四、把外部的承诺变成内部今天就能跑的测试
同一份公告里,Anthropic 重申了会「嵌入来自多个组织的独立第三方评估者」,并给他们提供与内部风险评估团队相当的流程、系统与数据访问权限;这些第三方将核查安全实践、报告事件、跟踪公告中列出的关键指标。对企业团队来说,这段话的实操价值不在 Anthropic,而在你自己:把它当作一份能力清单,逐条对比你现在能给出的审计接口。流程文档有没有?事件报告有没有?监控指标能不能导出?多数团队是在审计当天才发现缺哪一项,而不是提前。
# 4. Turn an external promise into an internal test you can run today
EMBEDDED_EVALUATOR_ACCESS = {
"processes": ["safety practice documentation", "incident reports"],
"systems": "comparable to what internal risk assessment teams already have",
"data": "the metrics described in the publication",
"deliverables": ["verify safety practices", "report incidents",
"monitor key metrics"],
"organisations": "multiple, independent of each other",
}
def gap_analysis(current_access, promised_access):
# Whatever the difference is, it is a to-do list you can close before an
# auditor asks for it. Most teams discover the gap during a review, not before.
return {k: promised_access[k] for k in promised_access
if promised_access.get(k) != current_access.get(k)}覆盖率、延迟、升级率:看见、来得及、有人管
五、一份可以在内部直接复用的报告格式
如果你想把这件事落地,最省力的方式是照抄它的结构:报告周期、指标族、所用刻度、实测值、方法说明、已知局限,以及「一旦越过阈值就触发什么动作」。例如 AL4 占比超过 50% 就冻结模型辅助的研发流程,直到外部评估者复核完毕;覆盖率低于 100% 就先列出所有未被监控的动作类别,再谈发版。最后别忘了复核周期——Anthropic 把节奏交给季度复核,你也应该给这份报告一个到期日,否则它会退化成一张漂亮的幻灯片。
{
"internal_ai_rd_report": {
"period": "2026-08",
"metric_family": "ai_led_rd",
"scale": "AL0-AL5",
"reported": {
"claude_leads_share": 0.26,
"at_or_above_collaborates_share": 0.92,
"fully_autonomous_share": 0.0
},
"method": "task catalogue + automation rating + aggregation, published",
"known_limitations": ["self-evaluation bias", "no cross-lab standard yet"],
"threshold_actions": {
"if_al4_share_gt_0_5": "freeze model-assisted R&D until an external evaluator reviews the pipeline",
"if_coverage_lt_1_0": "list every unmonitored action class before the next release"
},
"reviewed_by": "independent evaluator + platform engineering",
"review_due": "quarterly"
}
}📌 常见问题 FAQ
这三项测量具体是什么?
据 Anthropic 公告:一是 AI 自身完成了多少 AI 研发(AI 主导的研发比例),二是 AI 智能体的行为被监督与干预的程度,三是用于开发更强模型的算力如何分配。每一项都说明了测了什么、结果是什么、以及要做到可被他人定期核验还缺什么。
26% 是什么口径?
按 Epoch AI 的自动化等级(AL0–AL5),AL4「AI 主导」指给定高层提示即可端到端完成大部分任务、人类仅做监督。2026 年 8 月的快照显示 Claude 主导了 Anthropic 26% 的 AI 研发工作,且没有任何被测子集达到 AL5 完全自主。
为什么这些数字不能直接比较不同实验室?
Anthropic 明确列出两个障碍:缺少通用方法论;以及用自家模型评估自家系统会带来「裁判与被评估者共享失败模式」的风险。他们建议由独立第三方或其他开发者的模型核查,并设置限制竞争敏感数据交换的护栏。
智能体监督的覆盖率、复核延迟、升级率分别看什么?
覆盖率是动作前后经过监控的比例,回答「有没有被看见」;复核延迟是动作到被检查的时间(先自动、后人),回答「来不来得及」;升级率是被升级处理的活动占比,回答「看见了之后有没有人管」。
普通工程团队能从中抄走什么?
三件事:把自评数字连同方法与局限一起公开;把外部评估者的权限要求当成一份能力缺口清单;给每个指标设定阈值动作与复核到期日,让度量真正约束发布节奏。
🔧 推荐工具
📚 参考资料
- Anthropic (September 17, 2026) - Measurements for understanding the pace of AI development inside frontier labs: the three measurements, the Anthropic R&D Automation Index, Epoch AI's AL0-AL5 scale, Claude leading 26% of R&D work as of August 2026, and the plan to embed independent third-party evaluators
- Anthropic Newsroom - announcement index listing the September 17, 2026 measurement publication and the September 18, 2026 embedded-evaluation partnership
- NBC News - OpenAI and Anthropic scientists ask the U.S. government for tools to pace AI development (context on the "pace the frontier" proposal)