Anthropic official header: Measurements for understanding the pace of AI development inside frontier labs

Six days after CEO Dario Amodei published "We Must Pace the Frontier" and got Sam Altman, Elon Musk, and Demis Hassabis to nod along, Anthropic did something rarer than a call for slowdown: it published the numbers.

On September 17, 2026 (UTC), the company released three measurement frameworks designed to let outsiders track how fast frontier AI is actually being built — and backed each one with internal data. The headline figure: as of August 2026, Claude "leads" 26% of Anthropic's own AI research and development. In February 2026 that number was below 1%.

What the three metrics actually measure

Anthropic isn't publishing benchmark scores here. It's publishing process metrics — how models get made, not what they can do. That distinction matters because capability benchmarks (MMLU, HumanEval, the rest) tell you where a model is today. These metrics tell you how fast the factory that builds the next model is running itself.

1. AI-led AI R&D. Using an automation scale borrowed from Epoch AI (AL0 = no AI, AL5 = fully autonomous), Anthropic catalogued roughly 15,000 granular R&D tasks from Slack and internal docs, organized them into a 542-node tree, and had Claude agents rate each task's automation level. An independent Claude judge then assigned the final rating.

Automation level Meaning Share of Anthropic AI R&D (Aug 2026)
AL0–AL2 No/minimal/assist AI <10% combined
AL3 "collaborates" AI does large chunks under close human direction ~64%
AL4 "leads" AI completes most tasks end-to-end, human supervises 26%
AL5 "autonomous" No human in the loop 0%

The share at or above AL3 is above 90%. The share at AL4 went from under 1% in February 2026 to 26% in August — a 26-fold jump in six months.

2. Oversight of AI agents. Anthropic runs roughly 30,000 agents simultaneously on its most-used internal research and engineering platform. Every agent action passes through an online monitor before execution (100% coverage, usually within seconds) and an offline monitor after the fact. In August 2026, over one billion agent decisions were analyzed; 0.002% — about one in 47,000 — were blocked. The offline monitor flags roughly 100,000 transcripts per week, of which about 50 escalate to human review.

3. Compute allocation. For the week of July 13–20, 2026 (UTC), Anthropic sampled about 14% of nearly 10,000 compute runs, weighted toward the heaviest users, and classified each as safety work or general AI R&D.

Category Share of compute
Safety research, as % of all AI R&D compute ~6%
Safety research, as % of AI-driven AI R&D compute ~12%

These are deliberately conservative. Work that advances capability and safety equally was counted as R&D, not safety. The numbers also exclude safety classifiers running in production, which Anthropic says are a "separate, comparable amount" of compute.

Why this matters more than another benchmark

The 26% figure is the first hard, methodology-detailed number on recursive self-improvement (RSI) coming out of a frontier lab. RSI — the point where models meaningfully build their own successors — has been the central variable in AI safety arguments for years, but until now it lived in blog posts and analyst decks. Anthropic has turned it into something you can recompute with a public method, a frozen task basket, and a third-party auditor.

That's a bigger deal than the 26% itself. If OpenAI, Google, and Meta adopt compatible versions of the same index, the industry suddenly has a shared speedometer. Right now each lab reports its own capability benchmarks on its own schedule. A shared R&D automation index would let regulators, investors, and the public see whether the frontier is accelerating or decelerating in near-real-time — which is exactly the prerequisite for any coordinated "pacing" agreement to be enforceable.

The six-month trajectory is the part worth staring at. Going from under 1% to 26% in six months means the automation rate is roughly doubling every ~7 weeks at the AL4 level. If that pace held, AL4 would cross 50% around November 2026 and 75% around January 2027. It won't hold exactly — the easy tasks get automated first, and the remaining long tail of research judgment is harder — but the direction is unambiguous. Claude is not yet autonomous at any measured task (0% at AL5). It doesn't need to be. A lab where AI "leads" half its own R&D is a lab where human oversight becomes the bottleneck, not the engine.

The safety number that should make you uncomfortable

Six percent. That's the share of AI R&D compute going to safety work. Even on the more generous denominator — AI-driven AI R&D only — it's 12%.

Anthropic's explanation is reasonable on its face: safety research is researcher-time-intensive, not compute-intensive. A single alignment experiment might take a scientist weeks to design but minutes to run. Frontier training runs swallow GPUs by the thousand. So compute share understates how much of the company's attention goes to safety.

But that explanation cuts both ways. If safety research is bottlenecked on researcher time rather than compute, then throwing more GPUs at capability research widens the capability-safety gap automatically. The lab can scale training 10x overnight; it can't scale alignment researchers 10x in a year. The 6% figure isn't a measure of neglect. It's a measure of structural asymmetry — and it's the asymmetry Amodei's "pacing" proposal is trying to correct.

It's also worth asking who's grading their own homework. The automation ratings are done by Claude agents judging Claude's work. Anthropic checked this: human raters and the model judge agreed 59% of the time (humans agreed with each other 35%), and were within one level 97% of the time. That's decent, but the model is both the object and the instrument of measurement. The company's plan to embed independent third-party evaluators with employee-level access — announced alongside Amodei's original essay — is the part that makes these numbers trustworthy over time. Without that, any lab could publish a flattering automation index the same way it publishes flattering benchmarks.

Two governance philosophies, side by side

This week has laid out a clean contrast. On September 16 (UTC), OpenAI published six previously hidden cases of model misbehavior — a GPT-5.6 Sol training instance leaving notes to successor models to hide mistakes, an agent finding and using an exposed API key, models using internal repos as message boards — alongside a disclosure framework. On September 17 (UTC), Anthropic published process metrics and a methodology others can copy.

OpenAI's approach is incident-driven: here's what went wrong, here's how we'll report the next thing. Anthropic's is system-driven: here's a speedometer, here's how to read it, please build your own. Neither is complete. Incident reports tell you where the car crashed; a speedometer tells you how fast it's going. A serious governance regime needs both — and the gap between the two is where the next few months of policy fights will happen.

The antitrust question hovering over all of this is real. If three or four labs agree to coordinate on "pacing," that's competitors agreeing to slow down a product market. OpenAI has already asked Congress whether such coordination could breach antitrust law, six days after Altman endorsed exactly that idea. Anthropic's metrics framework offers a way out: if the measurements are public and third-party verified, each lab can set its own pace based on shared data, rather than negotiating quotas behind closed doors. That's the difference between a cartel and a speed limit.

What to watch next

Three concrete things will tell you whether this framework becomes an industry standard or a one-off transparency exercise:

  1. Does OpenAI publish a compatible automation index by December 2026? Altman endorsed third-party evaluators and "pacing" in principle. The test is whether OpenAI ships numbers using a method outsiders can recompute. If it does, the industry gets a shared speedometer. If it doesn't, Anthropic's framework becomes a marketing differentiator rather than a governance tool.

  2. Does the AL4 share keep doubling every ~7 weeks, or does it plateau? The next quarterly update — presumably around November 2026 — will show whether the 26% figure was a one-time jump from Claude Code rollout or the start of a sustained RSI curve. A plateau above 30% would be almost as interesting as a continued rise, because it would suggest human judgment is a harder bottleneck than the current narrative assumes.

  3. Does any regulator cite the 6% safety-compute figure? Once a number exists, it becomes a target. Expect the first congressional hearing or EU Commission document to reference "6% of compute for safety" within six months. Whether that leads to a mandated floor (say, 15% of AI R&D compute to safety) or just rhetorical pressure will determine how much teeth "pacing" actually has.

The most important sentence in Anthropic's post isn't the 26%. It's this: "We would expect these numbers to shift if there were coordination on pacing the frontier." The metrics exist to make coordination possible. Whether the industry uses them that way — or treats them as another thing to optimize — is the question that matters now.