Claude now does 26 percent of the work that builds the next Claude
Key takeaways
- Anthropic reported that as of August, Claude leads 26 percent of the work involved in building the next Claude, up from under 1 percent in February.
- More than 90 percent of that work has the model acting as a collaborator or better, rather than as a tool being operated step by step.
- The definition of leading the work is Anthropic's own and cannot be independently audited.
Under 1 percent in February. 26 percent in August.
That is Anthropic's own measure of how much of the work of building the next Claude is now led by the current Claude. Six months separates those two figures. Whatever shape you thought the curve had, that is the shape it has in one company's internal numbers.
The second figure matters more than the first. More than 90 percent of that 26 percent involves the model acting as a collaborator or better, rather than as a tool being operated step by step by an engineer. The distinction is the difference between autocomplete and a colleague who takes a task away and comes back with it done.
The caveat worth stating plainly
"Leads the work" is Anthropic's definition, measured by Anthropic, published by Anthropic. Nobody outside the company can audit what counts as leading, what counts as the work, or where the line between collaborator and tool was drawn.
That is not a reason to dismiss the number. It is a reason to treat it as directional rather than precise. The direction is consistent with what every frontier lab says quietly when nobody is writing it down: the single largest internal use of a frontier model right now is building the next frontier model.
Why this changes the shape of the business
If model development is increasingly done by models, the binding constraint moves. Research headcount stops being the thing you cannot buy quickly, and compute plus evaluation capacity becomes it instead.
That is a different business to run. It rewards whoever can secure accelerator supply and build enough evaluation infrastructure to check the output, rather than whoever can hire the most researchers. It also raises the cost of being wrong, because an error in a model that writes the next model propagates in a way an error in a person's code review does not.
The competitive picture reflects it already. When four frontier models land in a single week, the release cadence is not being set by how fast humans can write and test. Pricing pressure between the top models and the pace of iteration at Google point the same way.
What to watch
The number to track is not 26 percent. It is whether the next published figure is 40 percent or 30 percent, because the gap between those two tells you whether this is an exponential or a curve flattening against something nobody has named yet.
Anthropic has an obvious incentive to publish the figure again if it goes up, and a fairly obvious incentive not to if it stalls. Silence on it in six months would itself be an answer.