Anthropic says Claude now leads 26% of its model R&D, and asks rivals to publish the same numbers
On September 17, 2026, Anthropic published an announcement built around three sets of numbers, turning a conversation that had mostly happened in private into something comparable. The first: Claude now leads 26% of the company's model research and development. The second: roughly 90% of its R&D is done in collaboration with Claude. The third: in February of this year Claude led none of that work, and by August it led a quarter. Taken together, the three are more concrete than any judgement about how far away general intelligence is, because they describe how the division of labour inside a lab has shifted, not how many points a benchmark gained.
Anthropic splits the 26% figure into two carefully separated buckets. Lead means the model can take a high-level prompt and complete most of a given task end-to-end, but still under human supervision; the model cannot yet work completely autonomously. The remaining roughly 90% is collaboration, which Anthropic describes as the model doing large chunks of work under close human direction. The difference between the two buckets is not the model's ceiling but whether a human is still inside the loop. The fact that lead went from zero to 26% in six months, stated in the same breath as an admission that the model is not yet autonomous, is not comfortable reading: a lab has spent six months handing over a quarter of what used to be human work.
The timing of the announcement matters more than the numbers. Just over a week earlier, an Anthropic researcher resigned with a dire warning about the threats the technology poses to humanity, touching off a fresh public debate about whether frontier model development should be slowed down. Anthropic chief executive Dario Amodei has been one of the loudest voices calling for a slowdown. That gives the disclosure a subtle tension: the company argues for restraint at the industry level while reporting, honestly, how far its own internal automation has gone. Anthropic's proposed answer, in the accompanying blog post, is transparency. Its wording is that we should do everything possible to minimise the gap between what frontier labs know and what the public knows, by better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information. In other words, if the slowdown will not be agreed, at least let outsiders see.
Two numbers in the announcement are easier to skim past. The first is scale: about 30,000 agents were doing research and engineering work inside the company as of August. That order of magnitude explains why Anthropic devotes space to oversight. Only when there are tens of thousands of things to oversee does the rate at which monitoring catches misbehaviour become statistically meaningful, and however low the per-agent error rate is, multiplied by thirty thousand it becomes a daily event. The second is a commitment to embed external third-party evaluators inside the company to monitor safety efforts. Putting evaluators inside the building is a much heavier and harder arrangement than having them write reports from the outside, but it answers a specific question: once models start helping develop the next model, who decides that the curve is moving too fast.
What the disclosure does not answer may matter as much as what it does. It does not say how close Anthropic believes it is to recursive self-improvement, and it does not offer a single agreed standard for judging. What it does offer is a methodology others can copy: define precisely what lead and collaborate mean, publish on a fixed cadence, and make numbers comparable across labs. That is why the announcement closes by urging peers to publish the same metrics. Read it with one caveat: these are self-reported internal figures with no independent channel to verify them, and the line between lead and collaborate is drawn by the company itself, leaving room for interpretation. Treat it as a statement that the books are now being kept, more than as a report card.
🤔 Frequently Asked Questions
What exactly does the 26% mean?
It means Claude leads 26% of the company's model research and development work. Under Anthropic's definition, lead means the model can take a high-level prompt and complete most of a task end-to-end, but still under human supervision; it does not yet work fully autonomously.
How is the 90% collaboration figure different from the 26% lead figure?
The 26% is the share the model leads; the roughly 90% is the total share it takes part in, which Anthropic calls collaboration and describes as the model doing large chunks of work under close human direction. They are different measures and should be neither added together nor used interchangeably.
Where does the curve start?
In February 2026 Claude led none of that work; six months later, in August, it led a quarter. Anthropic did not say how close it believes it is to recursive self-improvement, only that models accelerating their own development could make it more challenging for humans to understand or control these systems.
What is Anthropic asking other labs to do?
To share similar metrics on a regular basis using a public methodology so the numbers can be compared over time and potentially across labs. Anthropic argues this helps outsiders judge how close frontier labs are to recursive self-improvement, and it committed to embedding external third-party evaluators inside the company to monitor safety efforts.
🛠️ Recommended Tools
- AI Token CounterThis story is about turning a share of work into a percentage. On any usage-billed model API the bill lands on tokens, so counting what a prompt and a reply each consume is the first step before talking about cost.
- Word CounterThe 26% and the 90% are two differently defined shares and are easy to mix up when citing. Counting words and characters per section while writing is cheaper than fixing the framing afterwards.
- Markdown EditorTurning the lab's original definitions and figures into structured notes is more reliable than leaving them scattered across browser tabs, and it makes it easy to check months later which bucket a given release used.
Summary
On September 17, 2026, Anthropic disclosed in an official announcement that Claude leads 26% of its model research and development, that the model is not yet fully autonomous and remains under human supervision, and that roughly 90% of its R&D is done in collaboration with Claude, meaning the model takes on large chunks of work under close human direction. The curve starts at zero in February 2026 and reaches a quarter in August. About 30,000 agents were doing research and engineering work inside the company as of August. Anthropic says models accelerating their own development could make it more challenging for humans to understand or control these systems, argues for better measurement and public reporting to narrow the gap between what frontier labs know and what the public knows, and urges other developers to publish comparable metrics on a regular basis using a public methodology. It also committed to embedding external third-party evaluators inside the company. One caveat remains: these are self-reported internal figures with no independent verification channel, the boundary between lead and collaborate is drawn by the company, and the announcement does not say how close Anthropic believes it is to recursive self-improvement. Every figure and claim here comes from AP reporting and Anthropic's own announcement, with no third-party speculation added.
Sources: AP News: Anthropic says its model Claude is helping to build the next version of itself
Anthropic 官方新闻页(公告原文所在渠道)