Anthropic's Claude 'Leads' 26% of R&D Tasks—With Supervision
New metrics reveal extensive AI involvement in model development, but human oversight remains mandatory and measurement gaps persist.

Anthropic measures AI's role in its own development
On September 17, 2026, Anthropic released three proposed metrics quantifying how much its AI model Claude contributes to research and development inside the company. The headline figure: Claude "leads" 26% of weighted R&D work and collaborates on more than 90%. Yet no category has reached full autonomy, and the measurements reveal significant structural limitations.
According to the Associated Press, the 26% figure represents work Claude completes largely from high-level prompts while a human remains in the loop. The distinction matters. "Leads" sits one tier below autonomous operation on Anthropic's five-level scale, which runs from no AI involvement through assistance, collaboration, and leadership to fully autonomous work.
The metric measures responsibility for established tasks, not control over research strategy, validation of scientific conclusions, or decisions about how future models should be built. Claude does not select Anthropic's priorities or initiate research programs independently.
How the denominator was constructed
Anthropic built its baseline from a random weekly sample of staff across model-development departments during a single month. A Claude research agent reviewed work records and documentation, extracted granular tasks, and organized them into a frozen hierarchy used for subsequent measurements.
The weighting system uses person-time as a proxy for importance. Each sampled employee receives equal weight for a week, divided evenly among their recorded tasks. The method does not measure actual hours spent, the value of outcomes, or the strategic importance of different projects.
This means the 26% figure represents the weighted share of baseline workload assigned the "leads" rating—not the share of Anthropic's ideas, decisions, experiments, or successful results produced by Claude. The frozen basket enables consistent tracking over time but cannot automatically capture new research categories, shifts in project importance, or tasks that disappear as methods evolve.
Claude evaluates its own contributions
For each category, a Claude agent reconstructs how work is performed and assigns an automation level. A separate Claude judge then rates the category, while human work-area owners provide comparison ratings without seeing the model-generated evidence or decisions.
This structure creates a dependency on Anthropic's own models throughout the evaluation pipeline. A judge from the same model family could share errors or blind spots with the systems being assessed, and the boundary between collaboration and leadership remains a judgment call rather than a directly observable threshold.
Oversight and safety metrics show reach, not accuracy
Quartz reported approximately 30,000 concurrent research and engineering agents running on Anthropic's most-used internal platform in August, with monitoring before and after every action. The block rate across more than one billion online-monitor decisions was 0.002%—roughly one in 47,000.
That intervention rate reveals how often monitors acted, not how many harmful actions they missed or how many blocked actions were actually harmless. Estimating detection accuracy would require independently labeled examples of acceptable and dangerous behavior, including failures that passed through undetected.
Safety work consumed about 6% of all AI R&D compute and 12% of AI-driven AI R&D compute during a one-week July snapshot. The different percentages reflect different denominators. Compute remains an incomplete proxy for organizational effort, and a single week cannot establish whether safety allocation is rising, falling, or stable.
Why it matters
These measurements offer rare transparency into how AI companies use their own models internally, but the gaps matter as much as the numbers. The 26% figure describes supervised task execution within a fixed framework, not autonomous research capability. Without repeated measurements under stable definitions and external validation of automation ratings and monitoring effectiveness, the metrics function best as an internal time series rather than proof of independent AI research capacity. For regulators and competitors watching the pace of AI-driven AI development, understanding what these numbers do not measure is essential.
These details were first reported by Automation Watch, with additional coverage from the Associated Press and Quartz.
This is an original analysis by the Omega editorial team. Source reporting: Automation Watch.
Want systems like this working for your business?
Book a Call
