Anthropic Shows Its Working
Claude "leads" 26% of Anthropic's own AI research and development as of August 2026, with more than 90% of the work at "collaborates" or above and nothing yet running fully autonomously. That figure sits in three measurements the Anthropic Institute has just published, offered to the public as a way to watch the loop where AI builds AI while the industry argues about pacing the frontier. The methodology sits in an appendix longer than the argument.
The ratings use a six-level scale borrowed from Epoch AI, assigned by a Claude judge reading evidence a Claude research agent gathered. Everything turns on where "collaborates" ends and "leads" begins, and that boundary is soft. Anthropic checked the judge against the staff who own each area: it matched them exactly 59% of the time, and they matched each other 35% of the time. The judge is more consistent than the people, which is the more interesting result, and it means the scale isn't measuring something its raters can reliably see. Ratings landed within one level 97% of the time, and one level is the whole distance between 26% and a rather different number.
The weighting deserves a look too. Each sampled person contributes one unit per week, split evenly across every task they touched, so a morning on a postmortem counts for as much as three days of pretraining work. That is a headcount proxy rather than a time proxy, and Anthropic calls it crude. The bigger question is where the basket came from: a Claude agent read the Slack messages and internal documents of a 20% staff sample to list what people worked on. Work that no human touched or discussed in July 2026 never entered the catalogue, and so cannot be rated as automated.
Compute is where the hedging piles up. Six percent of AI R&D compute went to safety over the sampled week, twelve percent of the AI-driven portion. Before you can react, you are told compute is an imperfect proxy, safety research is inherently compute-light, the estimate is deliberately conservative, efficiency gains shrink the share, and one week is not a trend. Each of those is true. The line that would settle the matter, that an independent third party could re-run the classifier on a random subsample, is written in the future tense.
Oversight is the strongest section, partly because METR has already red-teamed the offline monitor from outside. Roughly 30,000 agents, every action through a monitor before execution, one block in 47,000 across more than a billion August decisions. That is somewhere around twenty thousand blocked actions in the month, and Anthropic says humans review any blocked action within a week. It doesn't say how.
The contribution here isn't the numbers. It's the header repeated after each one: what any AI developer could report today. Outside researchers had already proposed metrics like these; the difference is that Anthropic ran them on itself and published the working. A method with its weak points attached is harder to walk back than a principle, and it hands outsiders something to pull on. METR did that in July, re-deriving researcher uplift from Anthropic's own published code figures and noting that the methodology behind Anthropic's lower estimate wasn't public.
Third-party evaluators with internal access, which is what a credible brake would need, appear here as something Anthropic plans to embed. Until they are, this is a company measuring itself with its own models and publishing the result, which is worth more than nothing and less than an audit.
Sources:
-
Measurements for understanding the pace of AI development inside frontier labs — Anthropic Institute
-
Toward an O*NET for AI R&D — Epoch AI
-
We Must Pace the Frontier — Dario Amodei
-
Because 8 ≈ e², Anthropic's researcher uplift is plausibly >2x — METR
-
Measuring AI R&D Automation — Chan et al., arXiv
Filed under AI & machine learning
This post is timestamped using Blockchain technology. Verify