26% of R&D: Anthropic Opens the Lab Door With Three Metrics

·11 min read·Evergreen Tools Team

On September 17, 2026, Anthropic published something unusual: not a model, not a benchmark score, but a proposal for measuring how fast AI development is moving inside a frontier lab. Three measurements - how much of AI R&D is performed by AI itself, how well the actions of AI agents are overseen, and how compute is allocated - each shipped with a snapshot from inside Anthropic. The value here is not the size of the numbers. It is that the question of what happens inside a lab has been rephrased as something a third party could, in principle, check.

Turning the inside of a lab into questions a third party can check

Turning the inside of a lab into questions a third party can check

1. The Scale That Makes the Percentage Readable

Anthropic borrows Epoch AI's Automation Level, or AL, which runs from AL0 (no AI involvement) to AL5 (AI operates fully autonomously, with no human in the loop). The middle levels carry the meaning. At AL3, AI collaborates: it can do large chunks of work under close human direction. At AL4, AI leads: it can complete most of the task end-to-end from a high-level prompt while the human supervises. As of August 2026, Claude was not operating fully autonomously for any measured subset of AI R&D work; it leads 26% of Anthropic's AI R&D work; and the share of work at or above AI collaborates is above 90%. Those three figures together are what make the 26% legible: not that AI does a quarter of the work, but that for a quarter of the work nobody has to specify each step.

# 1. The scale that makes the number readable: Epoch AI's Automation Level
AUTOMATION_LEVELS = {
    "AL0": "no AI involvement",
    "AL3": "AI collaborates: large chunks of work under close human direction",
    "AL4": "AI leads: most of the task end-to-end from a high-level prompt, human supervises",
    "AL5": "AI operates fully autonomously, no human in the loop",
}

# Anthropic R&D Automation Index methodology, in one sentence:
# catalogue every kind of AI R&D work, rate how automated each task is, aggregate.
SNAPSHOT_AUG_2026 = {
    "fully_autonomous_subsets": 0,      # Claude is at AL5 for no measured subset
    "claude_leads": 0.26,               # AL4
    "at_or_above_collaborates": ">0.90",  # AL3+
}

def progress_is_news(prev, cur):
    return "automation moved a level, not a percentage point" if prev != cur else "flat"

2. Why a Single Percentage Is Not a Metric

The report is candid about its own limits, which is the unusual part. First, there is no common methodology, so numbers are not yet comparable across labs. Second, and more interesting: they are using their own models to evaluate their own systems, which means the judge model could make the same kinds of errors as the model it is checking. Anthropic proposes three mitigations - verification by an independent third party, verification by other developers' models with guardrails limiting the exchange of competitively sensitive data, and using the measures as a trigger for stronger requirements, such as a fixed testing window before a new model is used for further AI R&D. The transferable rule for anyone doing internal AI governance: if the number cannot be verified, publish the method.

// 2. Why a single percentage is not a metric - the two stated obstacles
const reportingGaps = {
  methodology: "no common methodology across labs, so numbers are not comparable yet",
  judgeBias: "labs evaluate their own systems with their own models, so the judge "
             + "can share the same failure modes as the model under review",
};

const mitigationsAnthropicProposes = [
  "verification by an independent third party",
  "verification by other developers' models, with guardrails limiting the exchange of competitively sensitive data",
  "using the measures as a trigger for stronger requirements, such as a fixed testing window before a new model is used for further AI R&D",
];

// Design takeaway: if you cannot verify the number, publish the method.
If the number cannot be verified, publish the method

If the number cannot be verified, publish the method

3. Agent Oversight Is Three Numbers, Not One

The second measurement targets oversight of agents on Anthropic's own systems, and it decomposes into three distinct quantities: coverage, the share of an agent's actions that pass through a monitor before or after execution; review latency, the time between an action and its review, first by an automated monitor and then by a human; and escalation rate, the share of agent activities that are escalated. Steal the decomposition. Coverage answers whether anything is seen. Latency answers whether it is seen in time. Escalation rate answers whether anyone acts on it. Anthropic's stated motivation is the sharpest line in the piece: in monitoring data to date individual agents rarely misbehave, but when there are millions or billions of agents operating in the economy, even rare events can happen regularly.

# 3. Oversight of agents is three different numbers, not one
def oversight_metrics(agent_actions):
    return {
        # share of an agent's actions that pass through a monitor, before or after execution
        "coverage": len([a for a in agent_actions if a.monitored]) / max(len(agent_actions), 1),

        # time between an action and its review: first automated, then human
        "review_latency": {
            "automated_monitor_seconds": percentile(agent_actions, "auto_review_latency", 50),
            "human_triage_seconds":     percentile(agent_actions, "human_review_latency", 50),
        },

        # share of agent activity that escalates
        "escalation_rate": len([a for a in agent_actions if a.escalated]) / max(len(agent_actions), 1),
    }

# Anthropic's framing is the part worth stealing: individual agents rarely
# misbehave, but at millions or billions of agents, rare events happen regularly.

4. Turn a Public Promise Into a Test You Can Run Today

Alongside the metrics, Anthropic restates its plan to embed independent third-party evaluators from multiple organisations and give them access to internal processes, systems and data comparable to what internal risk assessment teams have, so those parties can verify safety practices, report incidents and monitor the key metrics. For your own team the useful move is not to read it as news about Anthropic but as a capability checklist for you. Do you have written safety-practice documentation an outsider could read? Incident reports with a defined format? Metrics that export? Most teams discover which of these are missing during a review rather than before one.

# 4. Turn an external promise into an internal test you can run today
EMBEDDED_EVALUATOR_ACCESS = {
    "processes": ["safety practice documentation", "incident reports"],
    "systems":   "comparable to what internal risk assessment teams already have",
    "data":      "the metrics described in the publication",
    "deliverables": ["verify safety practices", "report incidents",
                     "monitor key metrics"],
    "organisations": "multiple, independent of each other",
}

def gap_analysis(current_access, promised_access):
    # Whatever the difference is, it is a to-do list you can close before an
    # auditor asks for it. Most teams discover the gap during a review, not before.
    return {k: promised_access[k] for k in promised_access
            if promised_access.get(k) != current_access.get(k)}
Coverage, latency, escalation: seen, in time, acted on

Coverage, latency, escalation: seen, in time, acted on

5. A Report Format You Can Reuse Internally

The cheapest way to act on this is to copy the structure: a reporting period, a metric family, the scale used, the measured values, a statement of method, the known limitations, and a threshold action for each metric. If the AL4 share crosses 50%, freeze model-assisted R&D until an external evaluator reviews the pipeline. If coverage is below 100%, enumerate every unmonitored class of action before the next release. Then set a review date. Anthropic's version is a quarterly cycle; yours should have an expiry too, or the report degrades into a slide that is admired once and never revisited.

{
  "internal_ai_rd_report": {
    "period": "2026-08",
    "metric_family": "ai_led_rd",
    "scale": "AL0-AL5",
    "reported": {
      "claude_leads_share": 0.26,
      "at_or_above_collaborates_share": 0.92,
      "fully_autonomous_share": 0.0
    },
    "method": "task catalogue + automation rating + aggregation, published",
    "known_limitations": ["self-evaluation bias", "no cross-lab standard yet"],
    "threshold_actions": {
      "if_al4_share_gt_0_5": "freeze model-assisted R&D until an external evaluator reviews the pipeline",
      "if_coverage_lt_1_0": "list every unmonitored action class before the next release"
    },
    "reviewed_by": "independent evaluator + platform engineering",
    "review_due": "quarterly"
  }
}

📌 Frequently Asked Questions

What are the three measurements?

Per Anthropic's publication: the extent to which AI is building the next version of itself rather than being built by humans; the ability to oversee and intervene in actions AI agents take on internal systems; and the resources that power the development of more capable models. Each entry states what was measured, what the measurement showed, and what it would take to publish it regularly in a verifiable form.

What does the 26% figure mean?

On Epoch AI's Automation Level scale, AL4 means AI leads: it completes most of a task end-to-end from a high-level prompt while a human supervises. As of August 2026 Claude leads 26% of Anthropic's AI R&D work, with no measured subset at AL5 full autonomy.

Why are the numbers not comparable across labs?

Anthropic names two obstacles: the lack of a common methodology, and the self-evaluation problem - using your own models to evaluate your systems means the judge may share the errors of the model it is checking. The proposed fix is third-party verification or verification by other developers' models under data-sharing guardrails.

What do coverage, review latency and escalation rate each tell you?

Coverage is the share of an agent's actions that pass through a monitor before or after execution. Review latency is the time from action to review, automated first and human second. Escalation rate is the share of agent activity that gets escalated. Seen, seen in time, and acted upon.

What can a normal engineering team take from this?

Three habits: publish the method and the limitations next to any self-reported number; treat an external evaluator's access requirements as a gap list you can close early; and give every metric a threshold action and a review date so measurement actually constrains the release schedule.