In February 2026, Claude led under 1% of the AI research and development work happening inside Anthropic. By August, it led 26%. On the same internal platform, roughly 30,000 agents were running at any one time. Those numbers came from Anthropic itself, and they travelled fast — usually rewritten as some version of "AI now writes a quarter of the code at an AI company." That rewrite is wrong, and the way it is wrong is more interesting than the headline. The 26% is not a share of code. It is a share of tasks, weighted by human time, at a specific rung on a six-rung ladder. Understanding which rung changes what the number tells you about your own job.
What Anthropic Actually Measured — and What It Did Not
The source is a methodology page from the Anthropic Institute describing something it calls the Automation Index. The construction matters, because the construction is the caveat.
For each week of July 2026, Anthropic randomly sampled 20% of staff from every department in the model R&D loop. A Claude research agent read each sampled person's week through Slack and internal documentation and listed what they worked on. That produced roughly 15,000 granular tasks, which Claude then organised into a tree of 542 nodes — leaves with names like "eval platform defect diagnosis" and "RL sandbox egress." A separate Claude judge rated each node for how automated it was.
The scale is adopted from Epoch AI and runs from AL0 to AL5: no AI involvement, minimal involvement, AI assists, AI collaborates, AI leads, AI operates fully autonomously. The 26% is the weighted share of tasks sitting at AL4, where the model "can complete most of the task end-to-end from a high-level prompt, while the human supervises." Weighting is by person-time, so work that consumes more human hours counts for more.
Two details rarely survive the retelling. The first is that more than 90% of the measured work sits at AL3 or above — "AI collaborates" — which is a far larger number than 26% and a much weaker claim. The second is that the scope is the model R&D loop specifically, not all engineering at the company.
Read that definition again, because it contains a human. AL4 is not the machine working alone. Anthropic's own worked example is a broken nightly data pipeline: Claude finds the fault, fixes it, and writes up what happened — and an engineer still decides whether the fix ships. The number describes a division of labour, not a replacement. Anthropic states plainly that Claude "is not operating fully autonomously for any measured subset of AI R&D work." Nothing hit AL5.

The Measurement Problem Anthropic Disclosed Itself
Here is the detail almost no coverage mentioned, and it is in Anthropic's own write-up. The automation ratings depend on the judge model. When Anthropic checked the judge against humans, model and human agreed 59% of the time. When it checked humans against other humans, they agreed 35% of the time.
Sit with that second figure. Two experienced people looking at the same engineering task disagreed about how automated it was roughly two times in three. The model was not the unreliable narrator here — it was more consistent with humans than humans were with each other. But it does mean the underlying quantity is genuinely fuzzy. Ratings landed within one level of each other 97% of the time, so the picture is not noise. It does mean that a number like 26% carries a wider error bar than its two significant figures suggest, and that the line between "collaborates" and "leads" is a judgement call people make differently.
Anthropic flags a second limitation that matters more for forecasting: the task tree is frozen as of July 2026. The index measures how much of a fixed basket of work has been automated. If new kinds of R&D work are appearing — and at a frontier lab they almost certainly are — a frozen basket cannot see them. As Anthropic puts it, the measure "does not, on its own, tell us whether new kinds of work are appearing that humans have shifted onto." That single sentence undercuts the most popular reading of the chart, which is to draw the line from 1% to 26% and extend it to 100%.
Worth knowing about that line too: nobody measured 1% in February 2026 during February 2026. The index was built in July, and earlier months were reconstructed backwards by the same agent pipeline, restricted to evidence from the month being rated or earlier. That restriction is a careful design choice and prevents hindsight leaking in. It remains a retrospective reconstruction rather than a contemporaneous reading, which is a meaningful difference when the steepness of the curve is the whole story.
The Number You Are Thinking Of Is From a Different Post
If you had a figure in your head before reading this, it may well have been "80% of the code." That number is real, and it is not from the study above. It comes from a separate Anthropic Institute post, "When AI builds itself", published a day later: "As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude."
Different metric, different unit, different month. The 80% measures the share of lines merged to production attributable to Claude. The 26% measures person-time-weighted task categories at one rung of an automation ladder. Much of the coverage blended the two into a single "Anthropic said" narrative, which is how a reader ends up believing a quarter of the engineering happened without people while also believing four-fifths of the code did.
The same post reports that in Q2 2026 the typical engineer merged eight times as much code per day as in 2024 — and then does something worth crediting. It immediately undercuts its own statistic: "Lines of code is an imperfect measure, as it measures quantity over quality. So 8× lines of code/engineer/day in the second quarter of 2026 is almost certainly an overstatement of the true productivity gain."
A company publishing a flattering number and attaching "almost certainly an overstatement" to it is not the behaviour of a marketing document. Which points at where the distortion in this story actually happened. It was not manufactured at the source. It was manufactured in transmission, across dozens of near-identical articles that kept the figure and dropped the hedge.
30,000 Agents Is a Statement About Review, Not Headcount
The phrasing is precise: approximately 30,000 agents doing research and engineering work "at any one time" on Anthropic's most-used internal platform. Concurrent, not cumulative. One platform, not the whole company. Anthropic does not define what counts as a single agent, which is worth noting before anyone divides 30,000 by a headcount and reports a ratio.
What that scale really describes is a bottleneck moving. If thousands of processes are producing work end-to-end and no measured area is autonomous, then every one of those outputs terminates in a human decision. The constraint on the system stops being how fast people can write code and becomes how fast people can judge code they did not write.
That is not a speculative claim. The DORA 2025 State of DevOps report found AI adoption correlating with higher software delivery throughput and higher delivery instability at the same time — more change failures, more rework. DORA's framing is that AI acts as an amplifier: teams with strong testing, version control and feedback loops gain from it, while teams without them get the instability without the benefit.
The 2025 Stack Overflow Developer Survey shows the same pressure from the practitioner side, and its numbers are pointed. Adoption reached 84%, while trust in AI accuracy fell to 29% — down from 43% the year before. The most-cited frustration, named by 66% of developers, was "AI solutions that are almost right, but not quite." Another 45% said debugging AI-generated code was time-consuming.
Almost right, but not quite, is the expensive failure mode. Code that is obviously wrong gets discarded in seconds. Code that is plausibly wrong consumes a reviewer's full attention, and there is far more of it than there used to be. Generation got cheap. Verification did not, and the pipeline downstream of generation was not built for the new volume.

The Practitioner Question: Where Do Juniors Learn Now?
This is the part of the story that deserves more care than it usually gets, because it is where the data and the anecdote pull in different directions.
The traditional path into software engineering ran through work that was economically marginal but pedagogically dense. You fixed small bugs. You wrote tests for someone else's module. You updated documentation and, in doing so, learned where the bodies were buried. None of it was valuable enough to protect, which is exactly why juniors were allowed to do it — and it is the first category of work a capable model absorbs.
The labour data is consistent with that. Stanford Digital Economy Lab's "Canaries in the Coal Mine?" working paper, built on ADP payroll records running to June 2026, found that employment among workers aged 22 to 25 in highly AI-exposed occupations sat about 19% below where it would be had it kept pace with similarly aged workers in less-exposed occupations — a gap that widened from 15% a year earlier. Software development sits inside that exposed group, though the authors do not publish a developer-specific figure, so the 19% describes the category rather than the job. They are also explicit that these are descriptive patterns, not causal estimates.
The mechanism is the detail that matters most: the divergence runs through reduced hiring, not through people being let go. Firms are not firing juniors. They have stopped opening the door. Industry data points the same way — SignalFire's 2026 talent report puts new-graduate hiring at large tech firms down roughly 65% against a 2019 baseline, and down about 76% at early-stage startups.
The same research contains a finding that complicates any clean doom narrative. Declines concentrate in occupations where AI substitutes for human tasks. Where AI complements the worker, employment is flat or rising — particularly for experienced people. Whether software engineering is a substitution case or a complementation case is not settled by the technology. It is settled by how a given team chooses to deploy it, which means it is a management decision wearing a technological costume.
I want to be honest about the limits of what I can tell you here. I have not interviewed engineers inside Anthropic, and I am not going to invent a quote to make a section land. What I can offer is an inference with its reasoning exposed, which you are free to reject: if the entry-level tasks disappear but the need for senior judgement grows, the industry has quietly removed the bottom rungs of a ladder while advertising more jobs at the top. That arrangement works for exactly one cohort — the people already above the gap.
The Productivity Claim Deserves More Scepticism Than It Gets
There is a result every engineering leader quoting automation statistics should have to read first. In 2025, METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real tasks in their own large repositories. Tasks were randomly assigned to allow or forbid AI tools.
Before starting, the developers forecast AI would speed them up 24%. Afterwards, they estimated it had sped them up 20%. Measured, they took 19% longer on the tasks where AI was allowed. The striking finding is not the slowdown, which is one study on a specific population with early-2025 tooling and should not be over-generalised. It is the gap between experience and measurement: skilled professionals were wrong about their own productivity, in a flattering direction, by roughly 39 percentage points.
That should induce humility about self-reported automation gains everywhere, including in an index where a model rates how much work a model did. It is not evidence that Anthropic's number is inflated — the methodologies are entirely different, and Anthropic measured tasks rather than asking people how fast they felt. It is evidence that the direction of error in this domain is consistently toward optimism, and that the only reliable correction is measurement someone had a chance to fail.
It is also worth resisting the urge to stack these two findings against each other as though one refutes the other. METR studied veteran maintainers working on repositories they had known for years, with early-2025 tooling. Anthropic measured engineers on agent-native internal infrastructure with late-2026 models. Those are close to opposite conditions. The honest reading is not that one is right, but that the effect of these tools depends enormously on context — which is precisely why a number produced inside the most favourable context in the industry travels so badly.

Demonstration or Warning? Read the Denominator
The obvious corporate reading is that Anthropic has shown the way and every engineering organisation should now chase its own 26%. I think that reading is a trap, for a reason that has nothing to do with whether the technology works.
Anthropic is the most favourable possible environment for this result. It has unmetered access to frontier models, staff who build them, internal tooling built around them, and a codebase whose engineers are unusually equipped to review machine-generated work. The 26% was measured inside those conditions. A logistics company with a fifteen-year-old Java monolith and four contractors is not running the same experiment, and should not expect the same curve.
There is also a structural point worth stating plainly. Anthropic sells the model it is measuring. That does not make the measurement dishonest — the methodology page is unusually candid about its own weaknesses, more candid than most vendor research, and publishing your inter-rater agreement when it is 35% is not what a marketing document does. But when a company publishes a number that is favourable to its product, the number deserves the same scrutiny you would apply to a benchmark in a press release, and the caveats deserve to travel as far as the headline. They usually do not.
If you lead an engineering team, the useful question is not "what is our 26%." It is: what happens to our review capacity if generation becomes free? DORA's instability finding suggests most organisations discover the answer the expensive way. The teams that benefit are the ones that treat verification, testing and code review as the scarce resource and invest there first. The rest simply relocate their bottleneck and call it velocity. We take that view seriously in our own software engineering practice, and it shapes how we think about AI governance frameworks that increasingly treat autonomous agents as a distinct risk category.
So What Is an Engineer For?
Here is the question I do not think anyone has settled, and I am not going to pretend otherwise at the end of an article.
Every number above points the same direction: the scarce human contribution is moving from producing solutions to specifying problems and judging results. Anthropic's own definition of "leads" has a human writing the high-level prompt and deciding what ships. Both ends of that sentence are judgement. The middle — the part that used to be the job — is what got automated.
But there is an uncomfortable objection to the tidy conclusion that engineers should simply become problem-definers, and it is worth stating in its strongest form rather than dismissing it. The ability to specify a problem well and to recognise a bad solution quickly is not a separate skill from writing code. It is a residue of having written a great deal of it, badly, and having been wrong in ways that cost something. If that is right, then judgement is not the thing that survives automation — it is the thing automation slowly starves, because the practice that produces it is exactly the work being removed. We may be harvesting a stock of expertise that was accumulated under conditions we are dismantling.
Or that is nostalgia dressed as analysis. Every generation of engineers learned on abstractions that the previous generation considered dangerously far from the metal, and the profession survived compilers, garbage collection, and Stack Overflow. Perhaps judgement gets rebuilt on new foundations, and the engineer who never hand-traced a segfault develops equivalent instincts about model behaviour and system boundaries instead.
I genuinely do not know which of those is true, and I am suspicious of anyone who claims certainty this early with this much measurement error. So the question to argue about is this: when the writing is delegated, is the value of an engineer the ability to define the problem — or was defining the problem always just the thing that people who had written a lot of code turned out to be good at? If it is the second, an industry that stops hiring juniors is not streamlining. It is eating its seed corn, one absorbed task at a time. If you are working through what this means for your own team, we are happy to think it through with you.






