A Lab Measured Its Own Machine

On September 17, Anthropic published the first public run of what it calls the R&D Automation Index, a scorecard the company built to track how much of its own AI research and development work is now performed by Claude rather than by its human staff. The headline number: Claude currently leads 26 percent of that work, up from under 1 percent in February 2026, and more than 90 percent of relevant tasks now sit at or above the level Anthropic labels AI collaborates. Roughly 30,000 AI agents are running research and engineering work inside the company at any given time, based on a one-week measurement window from July 13 to July 20.

The index sorts every task on a six-step Automation Level scale built with Epoch AI, from AL0, no AI involvement, up to AL5, where the AI operates fully autonomously with no human in the loop. Anthropic says no task has reached AL5 yet. The point of publishing the number at all, the company says, is that the public needs more information about how fast a lab can now improve its own AI using that same AI, since that pace is normally invisible from outside.

The Watching Is Thinner Than the Working

A quarter of Anthropic's research now runs through Claude with a human supervising rather than doing the work directly, and the company's own release pairs that number with how closely it says it is checking. An automated online monitor reviewing Claude's research decisions in real time blocks just 0.002 percent of them, roughly one in 47,000. A separate offline review process flags only one or two transcripts out of every thousand for a human to look at afterward.

MeasureAnthropic's reported figure
Claude-led share of AI R&D26 percent (from under 1 percent in Feb 2026)
Work at or above AI collaborates levelOver 90 percent
Online monitor block rate0.002 percent of decisions (about 1 in 47,000)
Offline transcripts flagged for human review1 to 2 per 1,000

None of this means the work is unsupervised. It means the supervision is statistical rather than constant, catching a small, sampled share of decisions rather than reading all of them, at the exact moment the company says the range of decisions Claude makes on its own is widening.

Six Steps, and Nobody Is on the Last One Yet

Anthropic's scale describes what each level of automation actually looks like inside a research task, and the gap between the levels the company reports and the top of the scale is the real story in the numbers.

LevelWhat it means
AL0No AI involvement
AL1AI assists in a minor way
AL2AI assists more substantially, human still leads
AL3AI collaborates, doing large chunks of work under close human direction
AL4AI leads, completing work end to end from a high-level prompt while a human supervises
AL5Fully autonomous, no human in the loop at all

Anthropic's own quote frames the risk plainly: as the work becomes increasingly automated, agents could make more consequential decisions, such as which research direction to pursue next, and those decisions need safeguards against harmful behavior from a single agent as well as an oversight regime to catch problems across all of them. The company is describing AL4, not AL5, but the distance between the two is a single step, not a fixed wall.

A Company Writing Its Own Exam

Anthropic is not just reporting a number here. It is proposing that governments use measures like this one to set transparency obligations, and to define triggers for stronger requirements, such as a fixed testing window before a new model is allowed to take on further AI R&D work itself. That is a real, useful idea. It is also a lab writing the first draft of the exam it expects to be given, using telemetry only it currently holds.

No EU AI Act provision today ties a systemic-risk GPAI obligation to a lab's own automation level or to how much of its safety compute goes toward watching its automated share of the work, currently 6 percent of total AI R&D compute by Anthropic's account, rising to 12 percent within the AI-driven slice of that work. Any European business now buying frontier-model access for its own product roadmap should read this less as an Anthropic story and more as a preview of the question a regulator, or a customer's procurement team, will eventually ask every AI vendor: how much of what you sell me was built by something you were not fully watching, and how do you know.