
Anthropic has published a figure that most technology companies would rather not measure, let alone disclose. In a document released through its newly formed Anthropic Institute, the company says its own Claude models now lead 26 per cent of the research and development work behind Anthropic’s next generation of models, up from under 1 per cent in February. Six months, in other words, took the company from a position where its AI systems barely touched their own engineering to one where more than a quarter of that engineering is, in Anthropic’s own words, completed “end to end from a high level prompt” under human supervision.
The figure sits inside a broader number that is arguably more significant. Anthropic says over 90 per cent of its model research and development now happens at or above what it calls the “collaborates” level, meaning Claude is doing substantial chunks of the work under close human direction even where it is not formally leading the task. As of August, the company had roughly 30,000 agents running simultaneously on its internal research platform, working through problems that once sat exclusively with human engineers: why a training run kept failing, why a model underperformed on a particular benchmark, how to restructure a piece of infrastructure that had become unwieldy.
Anthropic is currently the leading commercial user of its own products, and the model most people encounter today is Claude Opus 4.8, released in late May and still the company’s flagship for coding, agentic work and extended reasoning tasks. It is a faster, cheaper version of the frontier model Anthropic had a year earlier, and it is very likely the model doing a meaningful share of the 26 per cent of engineering work the company is now describing. The company has not named a specific successor model in this disclosure. What it has described is closer to a process than a product: a growing share of the labour that produces future Claude models is itself being carried out by Claude.
The company built a specific tool to produce this number, which it calls the R&D Automation Index. Rather than asking managers to estimate how automated their teams’ work had become, researchers sampled a fifth of Anthropic’s staff each week through July, mining Slack conversations and internal documentation for roughly 15,000 granular examples of actual research and engineering tasks. Claude itself was then set to organise those examples into a hierarchical taxonomy of 542 categories, 378 of them fine grained enough to sit at the bottom of the tree. That taxonomy was frozen once built, so that the same yardstick could be applied at each future measurement point without the categories quietly shifting underneath the numbers.
Each task in the frozen basket is then rated on a five point automation scale borrowed from the AI research group Epoch AI, running from AL0, meaning no meaningful AI involvement, through to AL4, meaning the AI leads the task end to end with a human checking the output rather than directing the process. A judge model, itself built on Claude, assigns the ratings based on the documented evidence for how each piece of work was actually done. It is this judge that produced the 26 per cent figure, and it is worth dwelling on what that means: the system reporting how independent Claude has become is, in large part, Claude reporting on itself.
Anthropic does not hide this. The document states plainly that when the automated ratings were checked against a sample scored independently by human evaluators, the two agreed exactly only 59 per cent of the time, although the AI and human ratings fell within one level of each other 97 per cent of the time. That is a meaningful gap between a headline number and the confidence a reader should place in its precision. A model grading its own progress toward autonomy, even imperfectly, is not the same evidentiary standard as an outside audit, and Anthropic’s own methodology notes make that distinction rather than obscuring it.
There are further limits worth setting out before anyone treats 26 per cent as a fixed fact about the state of AI research. The frozen task basket, drawn from a single month’s snapshot of work, cannot capture genuinely new categories of research that emerge after it was built, so the figure describes automation of existing work rather than the shifting shape of the work itself. Compute allocation, which Anthropic also reports, is an imperfect proxy for effort: the company found that 6 per cent of its overall AI research compute goes to safety work, rising to 12 per cent of the compute used specifically by AI systems doing AI research, but notes that safety research is inherently less compute hungry than the large training runs that dominate capability work, so a low percentage there does not necessarily mean safety is being starved of attention. Anthropic also says it counted work that helps capability and safety in equal measure as safety work only when there was genuine doubt, a choice that deliberately understates rather than flatters its own safety commitment on paper.
The oversight architecture sitting underneath all this is described in more granular terms than most companies offer. Online monitors review agent actions in something close to real time and can block a dangerous action before it executes; Anthropic reports a block rate of 0.002 per cent, or roughly one in every 47,000 decisions reviewed. A separate layer of offline monitors reviews transcripts after the fact, escalating around one or two sessions per thousand to a human, which works out to roughly 50 escalations a week across the whole research operation. Anthropic is careful to note that this system has only been running for months rather than years, which limits how much confidence anyone should place in its ability to catch rare or unusual failure modes at scale.
None of this happens in a vacuum. The disclosure lands amid a live and increasingly public argument inside the AI industry about the pace of development, with Anthropic’s own chief executive among those who have called for a more cautious approach, and after a researcher’s departure from the company earlier this year that was accompanied by public warnings about the risks of advanced AI systems. Anthropic frames its own transparency here as a contribution to that argument rather than a distraction from it, explicitly urging rival laboratories to publish comparable figures using a shared, public methodology, so that the pace of automation inside frontier labs can be compared across the industry rather than taken on trust from any single company’s press release.
The honest reading of this disclosure sits somewhere between the two headlines it could easily generate. It is not evidence that Claude has begun autonomously building its successor: Anthropic states clearly that the process remains under human direction, that the model does not operate independently of that supervision, and that reaching genuine recursive self improvement, in which a system fully autonomously builds the one that replaces it, remains a distinct and unreached threshold that the company built this entire measurement framework specifically to watch for. But it is also not nothing. A jump from effectively zero to more than a quarter of a company’s own engineering work inside two quarters is a real and fairly dramatic shift in how Anthropic’s staff spend their time, whatever the precision of the exact number attached to it.
What deserves more scrutiny than the 26 per cent itself is the instrument that produced it. A company measuring how far its AI has progressed toward independence, using that same AI to do the measuring, and getting an answer that happens to align with a five point jump every few months, is a setup that would draw obvious scepticism if the roles were reversed and a rival were doing the reporting. Anthropic’s decision to publish its own caveats, including the uncomfortable 59 per cent agreement rate between its AI judge and human raters, is a genuinely unusual piece of candour for an industry that more often leads with the flattering number and buries the caveat in a footnote, if it appears at all. That candour is worth crediting. It does not, however, substitute for the kind of independent verification that a claim about a company’s own systems building the next version of themselves will eventually require if the public is expected to take the direction of travel, rather than just the headline figure, at face value.
The following content has been published by Stockmark.IT. All information utilised in the creation of this communication has been gathered from publicly available sources that we consider reliable. Nevertheless, we cannot guarantee the accuracy or completeness of this communication.
This communication is intended solely for informational purposes and should not be construed as an offer, recommendation, solicitation, inducement, or invitation by or on behalf of the Company or any affiliates to engage in any investment activities. The opinions and views expressed by the authors are their own and do not necessarily reflect those of the Company, its affiliates, or any other third party.
The services and products mentioned in this communication may not be suitable for all recipients, by continuing to read this website and its content you agree to the terms of this disclaimer.