Blog

AI productivity: METR's 19 per cent is out of date

METR's follow-up study flips the sign. What the research actually shows in August 2026, and why both camps quote it wrong.

The most-quoted number on AI productivity is out of date. In 2025 METR measured, across 16 experienced open-source developers, 19 per cent more time needed with AI tooling. The same organisation’s follow-up study, running from August 2025 with 57 developers, measures 18 per cent less time for the same people. Both intervals are so wide that neither result is established.

Why the 19 per cent figure is quoted wrongly today

The 19 per cent comes from a randomised study by METR, February to June 2025: 16 experienced developers, 246 issues in their own repositories averaging over 22,000 stars. The tooling was mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet.

The finding was uncomfortable and therefore widely quoted: the developers expected 24 per cent speed-up beforehand, believed afterwards they had gained 20 per cent, and were measured to be 19 per cent slower.

METR continued the work. The update of 24 February 2026 reports a second study running from August 2025: 57 developers, 10 of them from the original study and 47 new, 143 repositories, over 800 tasks. All values in the same metric — time needed with AI against without, where a minus means less time:

GroupPoint estimateConfidence interval
Original study, Feb–June 2025+19 %, i.e. more time+2 % to +39 %
Same developers, second study−18 %, i.e. less time−38 % to +9 %
Newly recruited developers−4 %, i.e. less time−15 % to +9 %

The result has not weakened, it has changed sign. More important is the difference in the intervals: the original study’s excludes zero, both new ones include it. Neither slowdown nor speed-up is established.

METR distrusts its own result, and gives clean reasons

METR itself writes that the data does not work as a productivity measure. Verbatim, on selection bias: “Because of the selection effects in our experiment, our data is only very weak evidence for the size of this increase.” And on the central estimate: “we believe it is likely a bad proxy for the real productivity impact.”

The reason is the most interesting point in the whole debate. METR cut compensation from 150 to 50 dollars an hour. That changed who took part at all, and which tasks got submitted.

“30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI”, the report says. One participant, in their own words: “I avoid issues like AI can finish things in just 2 hours, but I have to spend 20 hours.”

That is exactly where measurability breaks. Filtering out the tasks where AI helps most measures systematically low. METR concedes that the results probably underestimate the true effect. On top of that, some participants ran several agents in parallel, which devalues the timing.

If you are writing a business case for AI tooling right now and notice you are missing the defensible number: that is not a failure of your research. Half an hour with somebody who has read the primary sources is cheaper at this point than another week of searching.

The counter-figure from the other camp is from 2022

The famous 55 per cent speed-up is a vendor number from 2022. GitHub Next published it on 7 September 2022, together with Microsoft’s Office of the Chief Economist. Ninety-five developers, randomly split into two groups, one single task: write an HTTP server in JavaScript.

The Copilot group took 1 hour 11 minutes, the control group 2 hours 41 minutes. The confidence interval runs from 21 to 89 per cent.

Three things about it are worth knowing before putting the number in a presentation. It comes from the maker of the tool being tested. It measures a self-contained laboratory task with no existing codebase, no review and no operations. And it is four years old, from before agentic tooling existed.

Who collected which number, and who paid for it

The comparison below sorts the six numbers that actually circulate in this debate by origin, funding and shelf life. All retrievals as of 20 August 2026.

NumberWho collected itWho paidSampleStatus in August 2026
55 % fasterGitHub Next, MicrosoftGitHub, i.e. Microsoft95 people, one lab taskVendor number from 2022, not valid for today’s tools
19 % slowerMETRnon-profit, no vendor16 people, 246 issuesSuperseded by its own follow-up
18 % / 4 % fasterMETRnon-profit, no vendor57 people, 143 repos, 800+ tasksCurrent, but the interval includes zero
AI acts as an amplifierDORAGoogle CloudState of DevOps 2025Latest state, no 2026 report
3.45 % vs 7.40 % breaking changesFerdous et al.academic, arXiv7,191 agent PRs, 1,402 human PRsCurrent
Context files ineffective, +20 % inference costGloaguen et al., ETH Zurichacademic, arXivSWE-bench plus real repositoriesCurrent, revised June 2026

The funding column is not a courtesy. DORA is a Google Cloud programme, and Google sells competing tools in Gemini, Jules and Antigravity. That does not make the methodology bad; it is rightly well regarded. But a report that helps companies defend AI budgets has a sender with an interest in AI budgets.

The amplifier finding is the most defensible statement in the research

The most stable insight is that AI tools amplify existing engineering maturity rather than replacing it. DORA puts it this way in the State of DevOps 2025: “AI acts as an amplifier, but the greatest returns come from focusing on the underlying sociotechnical systems.”

Two academic papers from 2026 show the same mechanism from below. The study “Safer Builders, Risky Maintainers” by Ferdous, Banik, Chowdhury and Shamim compared 7,191 agent-generated pull requests with 1,402 human ones from Python repositories and detected breaking changes through AST analysis.

ContextBreaking change rate
Code generation, agents3.45 %
Code generation, humans7.40 %
Agents on refactoring6.72 %
Agents on chore changes9.35 %

On new construction, agents break compatibility less often than humans. On maintenance work that inverts. The authors additionally describe a “confidence trap”: pull requests with a high self-reported confidence score from the agent contain breaking changes too. The score does not work as a review filter.

The second paper, “(Im)Paired Programming” by Balepur and colleagues, had 54 students build a website, with an agent or with a chatbot, and then tested comprehension through questions and through an extension task without an agent. Finding: agents help you finish and harm understanding, and it is precisely the comfortable interaction forms — copy-paste prompts, auto-accepted edits — that correlate with the worst understanding. The caveat belongs with it: 54 students are not an engineering team.

Why this supports the amplifier thesis: in both papers it is not the tool that decides but where and how it is used. The same observation from the practice side, without numbers, is in AI writes the code. Who writes the architecture?

What the range means for capacity planning

The second study’s confidence interval is too wide to budget with. Let us work it through in Swiss francs so the order of magnitude is visible.

A model calculation, not a survey: a team of ten developers, 1,800 working hours per person per year, so 18,000 hours. As an hourly rate we take CHF 160, the midpoint from the published worked example at gryps.ch for software engineering in German-speaking Switzerland (retrieved 20 August 2026). The annual capacity therefore corresponds to CHF 2,880,000.

Scenario from the METR intervalTime neededHours per yearValue
Lower bound−38 %6,840 hours freedCHF 1,094,400
Point estimate−18 %3,240 hours freedCHF 518,400
Upper bound+9 %1,620 hours extraCHF 259,200

Between the two bounds lies CHF 1,353,600. Nobody plans a budget on a range like that, and that is the research’s honest answer rather than its failure.

The calculation deliberately transfers METR’s values crudely onto an annual capacity. METR measured issue-shaped tasks in mature open-source repositories, not full working years with meetings, operations and incidents. Anyone transferring the numbers one to one onto their team makes the same mistake as the camps they are criticising.

Measure your own baseline instead

The most useful response to an unclear body of research is your own before-measurement. Two metrics are enough to start, and both already sit in your repository.

First, a proxy for change failure rate, out of the git log over 90 days:

# All commits in the last 90 days
git log --since="90 days ago" --no-merges --oneline | wc -l

# Of those, reverts and hotfixes
git log --since="90 days ago" --no-merges --oneline \
  --grep="^Revert" --grep="^hotfix" --grep="^fix!" -i | wc -l

Second, the size of changes per author, to separate agent-generated from human pull requests:

gh pr list --state merged --limit 200 \
  --json number,author,additions,deletions,mergedAt \
  --jq '.[] | [.author.login, .additions + .deletions] | @tsv'

Collect both a month before rollout and three months after. Those two numbers say more about your team than any study in this article. How to record the decisions behind them so the measurement is still interpretable in a year is in Documenting architecture decisions.

What does not exist in August 2026

Three frequently expected sources do not exist at the time of writing, and anyone citing them is citing something invented.

  • No State of DevOps report 2026. dora.dev/research lists 2025 as the most recent entry. The only new thing is an interim product, the “ROI of AI-assisted Software Development report”, last updated 22 April 2026, with an ROI calculator.
  • No Stack Overflow Developer Survey 2026. The archive lists 2011 to 2025. The 2025 edition remains current.
  • No Technology Radar Vol. 35. Vol. 34 of April 2026 is the current issue, with “Putting coding agents on a leash” as one of the four themes.

What we advise against

We advise against three things, and the first concerns the numbers in this article itself. First: building an ROI business case on a point estimate, from whichever camp. The defensible statement from the research is a range, not a value.

Second: using an agent’s confidence score as a review filter. The confidence trap from “Safer Builders, Risky Maintainers” is exactly that mistake, empirically demonstrated. Grade review depth by task type instead. Greenfield work can run more loosely; refactoring and maintenance need strict control.

Third: auto-accept as the default setting. According to “(Im)Paired Programming” it is the interaction form with the worst comprehension outcome.

Where our recommendation does not fit: throwaway prototypes and spikes. If you are checking feasibility for a week and deleting the code afterwards, you do not need graded review depth. The effort pays from the moment somebody has to read the code a year later.

And a Swiss detail that gets lost in the international debate: the methodologically cleanest work on context files for coding agents comes from ETH Zurich. Gloaguen, Mündler, Müller, Raychev and Vechev showed in February 2026, revised in June, that AGENTS.md files do not generally improve the success rate but raise inference costs by more than 20 per cent. Three of the authors are affiliated with LogicStar AG in Zurich; Martin Vechev leads the Secure, Reliable and Intelligent Systems Lab at ETH.

Frequently asked

Is it true that AI makes developers 19 per cent slower?

That was the finding of a METR study from February to June 2025 with 16 experienced open-source developers. The same organisation’s follow-up, running from August 2025 with 57 developers and over 800 tasks, measures the opposite: 18 per cent less time needed among the returning participants. Anyone offering the 19 per cent as evidence today is quoting a superseded result.

Where does the “55 per cent faster” figure come from?

From a study by GitHub Next and Microsoft’s Office of the Chief Economist, published 7 September 2022. Ninety-five developers wrote an HTTP server in JavaScript, with or without Copilot. It is a vendor number from 2022, measured on a single laboratory task with no existing codebase, no review and no operations.

Which statement from the research is the most defensible?

The amplifier finding. DORA states in the State of DevOps 2025 that AI acts as an amplifier and that the greatest returns come from work on the underlying sociotechnical systems. Two academic papers from 2026 show the same mechanism in detail: the effect depends on the task type and the way of working, not on the tool.

Why is the productivity effect so hard to measure?

Because the participants influence the measurement. When METR cut compensation from 150 to 50 dollars an hour, 30 to 50 per cent of developers said they were not submitting certain tasks because they did not want to do them without AI. Filtering out exactly the tasks where AI helps most measures systematically low.

What should you measure in your own team instead?

Two metrics from your own repository, once before and once three months after introduction: the rate of reverts and hotfixes over 90 days as a proxy for change failure rate, and the change size per author, to separate agent-generated from human pull requests. Both take a few minutes with git log and gh pr list.

Let’s put your numbers next to them

Bring your git history and the question your management asked you. We collect the baseline together and tell you which of the studies quoted here actually applies to your case and which does not.

Book a slot or write to us. How we support teams using these tools is under Artificial Intelligence & Machine Learning.

A conversation, not a newsletter

Let's talk about your system

If this article describes something you recognise, a conversation is the shortest route to an answer.

Let's talk