Do AI Coding Assistants Make Developers Faster? What the Measurements Actually Show

Four independently run studies, from a 2023 Upwork task to a 2026 replication METR itself distrusts, give four different answers — and the differences are the finding.

EduFabTech · 30 September 2026 · 8 min read · 1 views
A gauge needle hovers between "slower" and "faster," illustrating that four studies on AI coding assistants reached four different verdicts (+55.8%, −19%, +26.08%).
EduFabTech · Own work

In July 2025, a nonprofit research organization called METR published a randomized controlled trial with an inconvenient result: when sixteen experienced open-source developers were allowed to use current AI coding tools on real issues from their own repositories, they took 19% longer to finish than when they worked without them. The same developers, asked to predict the effect beforehand, guessed AI would cut their time by 24%. Asked again afterward, having just lived through the slowdown, they still guessed they had been sped up by 20%. The gap between what they felt and what a stopwatch recorded is, on its own, worth sitting with before any argument about code quality or tooling.

That result sits awkwardly next to a different, earlier number that circulates constantly in industry discussion: developers using GitHub Copilot completed a task 55.8% faster than a control group, according to a 2023 study run by Microsoft Research and GitHub economists. Both numbers are real measurements, published by credentialed researchers, with methods available for inspection. Neither is wrong. They measured different things, on different tasks, with different developers, at different points in a fast-moving technology's life. The purpose of this piece is not to declare a winner but to lay the four major measurements side by side, dated and attributed, so the actual shape of the evidence is visible.

This matters beyond curiosity about one product category. AI coding assistants are now embedded in the default toolchain of a large fraction of professional software development, and organizational decisions — hiring plans, tooling budgets, code-review policy — are being made on the back of productivity claims that rarely specify which study, which task, or which year they come from.

A diverging bar chart lines up all five measurements — from the 2023 Upwork task's +55.8% to the 2025 METR trial's −19% — against a "no change" baseline, with the discarded 2026 replication shown hatched.
A diverging bar chart lines up all five measurements — from the 2023 Upwork task's +55.8% to the 2025 METR trial's −19% — against a "no change" baseline, with the discarded 2026 replication shown hatched.EduFabTech · Own work

What the 2023 vendor-adjacent study actually measured

The most widely cited "AI makes developers faster" number comes from a controlled experiment described in a 2023 paper by Sida Peng, Eirini Kalliamvakou, Peter Cihon and Mert Demirer, published via arXiv and affiliated with Microsoft Research. The researchers recruited 95 professional programmers through the freelance platform Upwork and assigned a single, well-specified task: implement an HTTP server in JavaScript, as fast as possible, in a timed session. The group with access to GitHub Copilot finished in an average of 1 hour 11 minutes; the control group averaged 2 hours 41 minutes — a 55.8% difference.

The design is genuinely a randomized experiment, and the effect size is large enough that it is very unlikely to be noise. But the task was chosen precisely because it is the kind of task an autocomplete-style assistant is good at: a common, well-documented, greenfield component with a clear specification and no existing codebase to understand. It says something true and useful about a narrow slice of programming work. It does not, on its own terms, say anything about maintaining a five-year-old production codebase, negotiating an ambiguous bug report, or reviewing a large pull request — the paper does not claim otherwise, but the 55.8% figure is frequently quoted stripped of that context.

The 2025 randomized trial that found the opposite

METR — an organization that primarily evaluates frontier AI systems for capability and safety — ran a different kind of trial, published in July 2025 as "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity." Instead of a toy task, the researchers recruited sixteen developers who each had an average of five years' contribution history on large, mature open-source repositories (averaging over 22,000 GitHub stars). Each developer supplied a backlog of real issues — bug fixes, features, refactors — that they intended to complete regardless of the study. Each of the resulting 246 tasks was randomly assigned to an "AI allowed" or "AI disallowed" condition, with developers free to use whatever tools they wished, predominantly Cursor with Claude 3.5 and 3.7 Sonnet — the frontier models available in the February–June 2025 window the study covered.

The headline result was a 19% increase in completion time when AI use was permitted, with a wide confidence interval (roughly +2% to +39%) reflecting the small sample. The authors were explicit about what the result does and does not support: it is not evidence that AI tools fail to help developers in general, only that in this specific setting — high codebase familiarity, large and mature repositories, experienced maintainers, real production stakes — the tools slowed people down, largely because of time spent reviewing, correcting and integrating AI-generated suggestions rather than time saved writing code from scratch.

Field experiments inside real companies

A third data point comes from inside industry rather than a lab or a freelance marketplace. A paper by Zheyuan Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng and Tobias Salz, accepted at Management Science in 2026, combines three separate randomized controlled trials run inside the ordinary business operations of Microsoft, Accenture and an unnamed Fortune 100 company, covering 4,867 developers who were randomly given or denied access to GitHub Copilot. Pooled across the three firms, developers with Copilot access completed 26.08% more tasks (standard error 10.3%) than those without it, measured by weekly pull-request throughput. The gain was concentrated among less experienced developers, a pattern that recurs across several of these studies and cuts against a simple "AI helps everyone equally" reading.

This study sidesteps some of the criticisms of the 2023 Upwork experiment — it uses real production work at real companies rather than a single artificial task — but it measures a different outcome again: pull requests completed per week, not time-to-completion on a fixed task, and not code that necessarily required deep pre-existing familiarity with the surrounding system the way METR's open-source maintainers had.

What a 3,000-person cross-industry survey adds — and complicates

The 2024 Accelerate State of DevOps Report, produced by Google Cloud's DORA research program from a large annual practitioner survey, offers a fourth and messier picture. It found that for every 25-percentage-point increase in a team's reported AI adoption, individual-level productivity rose by an estimated 2.1% and job satisfaction rose 2.6% — small but positive. At the same time, the same 25-point increase in AI adoption was associated with a 1.5% decrease in software delivery throughput and a 7.2% decrease in delivery stability. Trust lagged further behind both: only about 24% of respondents reported trusting AI-generated code "a lot" or "a great deal," while roughly 39% reported little to no confidence in it.

AI adoption significantly increases individual productivity, flow, and job satisfaction — but it does not translate automatically into better team-level or organizational software delivery performance.

That combination — individuals feeling and reporting more productive, while system-level throughput and stability degrade — is not necessarily a contradiction. It is consistent with a world where AI assistance speeds up the writing of code while adding downstream review, integration and stabilization costs that show up in different metrics, collected from different people, on a different timescale. It is also drawn from a large annual practitioner survey rather than a controlled experiment, so causal claims from it should be read as associations, not effects.

Why METR abandoned its own follow-up

The most methodologically interesting development came in February 2026, when METR published a candid account of why its attempt to replicate and extend the original trial had failed on its own terms. Running a second study from August 2025 with 57 developers (10 returning, 47 new), the organization found a swing to an estimated 18% speedup among returning developers and a roughly 4% speedup among new ones — the opposite sign from the original result, but with confidence intervals wide enough to include zero and, for the returning-developer group, wide enough to remain consistent with the original slowdown.

METR did not present this as evidence that AI had "caught up." Instead, the team reported that the second study was compromised by selection effects severe enough to make the signal unusable: 30–50% of developers avoided submitting tasks they would not have wanted to complete without AI assistance, developers increasingly declined to participate at all if it meant working without AI tools, task types shifted toward work suited to AI's strengths, and a lower pay rate ($50/hour versus the original $150/hour) changed who volunteered. METR's own conclusion was that the study no longer measured what it was designed to measure, and the organization announced it is redesigning its methodology — shorter, better-compensated sessions, observational data, and fixed-task designs among the options under consideration.

A two-panel comparison shows the same METR developers predicting and later believing they were 20-24% faster with AI, set directly against the stopwatch's actual finding of 19% slower.
A two-panel comparison shows the same METR developers predicting and later believing they were 20-24% faster with AI, set directly against the stopwatch's actual finding of 19% slower.EduFabTech · Own work

Reading the four measurements together

Laid side by side with dates attached, the studies describe a technology whose effect depends heavily on what is being measured and who is doing the work, not a single number that moves cleanly over time:

StudyYearSettingResult
Peng, Kalliamvakou, Cihon, Demirer202395 Upwork freelancers, single greenfield HTTP-server task+55.8% faster with Copilot
Cui, Demirer, Jaffe, Musolff, Peng, Salz2026 (Management Science)4,867 developers, three companies, real production work+26.08% more tasks completed
DORA / Google Cloud2024Large annual cross-industry practitioner survey+2.1% individual productivity; −1.5% throughput; −7.2% stability per 25-point AI adoption rise
METR202516 experienced maintainers, real issues on mature open-source repos−19% (slower) with AI allowed

A pattern worth naming: the largest positive effects come from the tasks with the least pre-existing context — a fresh task with a clean specification, or, in the DORA data, an individual's felt sense of their own output. The negative or mixed effects appear where developers already carry deep familiarity with a large, mature system, where correctness and integration cost dominate, or where the outcome is measured at the level of a team or delivery pipeline rather than an individual task. That is a substantive finding in its own right, independent of any single headline number: the productivity effect of current AI coding tools is not a fixed quantity — it is a function of task type, codebase familiarity, and what is being measured, and no single study, including METR's, claims to generalize beyond its own setting.

What is still not known

None of the four measurements above answer, or claim to answer, what happens with the AI models and tooling available in late 2026 rather than the specific tool-and-model combinations tested — Copilot circa 2022, or Cursor with Claude 3.5/3.7 Sonnet in the February–June 2025 window. Capability and tool design both continue to change quickly enough that a study is, in a meaningful sense, dated the moment its data collection ends. METR's own July 2025 paper says this explicitly: it does not claim evidence about future AI systems, only about the ones its participants actually used. Any claim about "AI coding assistants" without a named model, a named study, and a stated year is not a measurement — it is an assertion standing in for one, and the honest response to encountering such a claim is to ask which study it is quietly borrowing credibility from, and whether that study's setting resembles the one the claim is being applied to.


References
  1. Joel Becker, Nate Rush, Beth Barnes, David Rein. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR, 2025. link
  2. METR. We are Changing our Developer Productivity Experiment Design. METR (blog), 2026. link
  3. Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv (Microsoft Research / GitHub), 2023. doi:10.48550/arXiv.2302.06590
  4. Zheyuan (Kevin) Cui, Mert Demirer, Sonia Jaffe, Leon Musolff, Sida Peng, Tobias Salz. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Management Science (INFORMS), 2026. doi:10.1287/mnsc.2025.00535
  5. DORA (DevOps Research and Assessment), Google Cloud. Accelerate State of DevOps Report 2024. Google Cloud / DORA, 2024. link
  6. Google Cloud. Announcing the 2024 DORA report. Google Cloud Blog, 2024. link