Contents
When a model dumps out a function in seconds, it feels obvious: developers should finish work faster. In practice, the measurements argue with each other. In one controlled experiment, the GitHub Copilot group finished a task about 55% sooner. In a randomized METR study on mature open-source projects, permission to use AI increased task time by about 19%. Both numbers can be true — if you know what was measured.
Key takeaways
Typing speed ≠ development speed. Draft generation is only one term in the sum: intent, verification, debugging, tests, integration, and review can erase the whole gain — or make the total longer.
On a short, formalized task, the speedup is often large. In the GitHub/Microsoft experiment, 95 developers built an HTTP server in JavaScript: with Copilot about 71 minutes on average, without about 161 minutes (roughly −55% time).
On real tasks in a large codebase, the effect can be negative. METR (early-2025 data) on 16 experienced open-source developers and 246 tasks recorded about +19% time with AI allowed — even though people expected roughly a 24% speedup.
Subjective productivity lies. After METR, participants still estimated they were about 20% faster, while the clock said the opposite. That is the productivity illusion.
Compare “human” and “human + AI,” not “human vs model.” The model usually acts as an amplifier: generator, pair programmer, review and docs helper. The 2026 question is not “will AI replace programmers,” but how much product one person can shepherd at the same decision quality.
The real question: what are we speeding up
Picture a factory. A press stamps a part in a fraction of a second. But a car leaves the line only after assembly, inspection, and finishing. Debating “press speed” helps the blanking shop and does nothing for the customer waiting for a finished car. Software is the same: line-typing speed is easy to confuse with delivery speed.
Separate ideas that AI conversations often glue into one:
- code-writing speed — how fast a draft appears;
- task-completion speed — time from start to “done” by the team’s criteria;
- code volume — lines or characters produced (a weak metric on its own);
- functional correctness — does it pass acceptance and tests;
- quality and maintainability — readability, fit with architecture, cost of future edits;
- time on review and fixes — review, rollbacks, debugging rabbit holes;
- total cost — person-hours, incidents, rework.
The article’s main claim is simple and hard: code-generation speed and software-product development speed are not the same thing. While you measure only the first, any “AI speedup” percentage will look more convincing than it is for release.
Why experiments beat feelings
After a week with an assistant, a developer often feels: “I write twice as fast.” That feeling is honest as a flow experience: less routine, less blank screen, more sense of progress. An objective check still has to answer other questions.
How long did the whole task take? Was the result correct? Did tests pass? How many edits after review? Can the change go to production? Does it fit the existing system — not only “compile on my machine”?
Hence a principle worth posting next to the dashboard: developer productivity ≠ lines of code per hour. Lines are a byproduct. The product is a working, checked, maintainable change. That is why we unpack controlled and field studies below, not only social-media testimonials.
Copilot on a formalized task: plus fifty-five percent
One of the most cited experiments is Microsoft Research and GitHub’s work around GitHub Copilot. In a controlled setup, 95 professional developers took part. They were randomly split into two groups, given the same JavaScript task — implement an HTTP server — and timed for duration and success.
The Copilot group finished in about 1 hour 11 minutes on average. The group without an assistant — about 2 hours 41 minutes. That is roughly a 55% speedup in time (GitHub publications use “55% faster”; the confidence interval was wide, but the direction of the effect was stable). Successful completion: about 78% versus 70%.
Sources: Microsoft Research, GitHub’s post on productivity and developer “happiness”.
The limitation is critical for interpretation. The task was relatively small, bounded, and well formalized: shared language, clear done criteria, automatic checks. That is closer to “stamping a part” than to “shipping a car.” You cannot automatically translate the result into “every developer now programs 55% faster” — especially in a monolith with a fifteen-year history, where the bottleneck is not typing an HTTP handler but understanding consequences.
Code quality: what GitHub measured
A separate question is not only “faster,” but “better.” In GitHub’s 2024 study with 202 developers (Python experience, random Copilot access, the same API-endpoint task, then automated tests and blind expert review) they looked at functionality, readability, reliability, maintainability, conciseness, and likelihood of code approval.
According to GitHub:
- the chance of passing all 10 unit tests was about 53.2% higher for the Copilot group;
- blind review found fewer readability issues;
- the chance of code approval was about 5% higher;
- small but statistically significant gains in readability, reliability, maintainability, and conciseness (on the order of a few percent).
Source: Does GitHub Copilot improve code quality?.
A conflict-of-interest caveat belongs here: the company whose product is evaluated runs the study. That does not void the methodology or make the numbers “fake,” but it obliges reading them next to independent work — including METR, where the picture is less rosy. Quality on a lab task with ten tests is also not the same as quality in a production incident six months later.
METR 2025: plus nineteen percent time on real tasks
METR’s study is especially interesting because it leaves the tutorial HTTP task behind. It is a randomized controlled study on 16 experienced open-source developers, 246 real tasks in large mature projects (millions of lines of code), with participants who had about five years of experience on average with those repositories. Tasks included fixes, features, and refactoring. Source: METR PDF.
Result: allowing AI increased task time by about 19%. Before the experiment, people expected roughly a 24% speedup. Afterward they subjectively estimated about a 20% speedup. Objective hours showed a slowdown. In METR’s 2026 update the wording is refined as “about a 20% slowdown” on early-to-mid-2025 data — the same order of magnitude.
This is not “AI is useless.” It is “in this regime, on these people and tasks, the end-to-end cycle got longer.” The gap between expectation and the stopwatch is the plot of the next section.
The productivity illusion
When the subjective estimate says “I got 20% faster” and the measurement says “the task took 19% longer,” a separate phenomenon appears: the productivity illusion. People feel progress because the screen is rarely empty, options appear quickly, and routine is delegated. But total time to an accepted change grows because of verification, integration, and debugging solutions that are foreign to their mental model.
The illusion is dangerous for management. A team can adopt a tool, report “speed growth” from surveys, and simultaneously lengthen the path to merge. A “lines per hour” dashboard and a “how satisfied are you” survey amplify the error. You need hours to readiness, defects after merge, and rework cost.
Why an experienced developer in a large codebase may slow down
Back to the factory. If you already know the assembly line by heart, a new machine that stamps “almost right” parts forces you to stop the conveyor more often and check tolerances. An experienced author in a familiar repository sits in exactly that seat.
Context. The model has to guess architecture, conventions, dependencies, legacy constraints, and unspoken rules that live in people’s heads. The more hidden context, the more often a draft is technically plausible and architecturally alien. More on the limits of understanding heritage — in the AI-on-a-15-year-project breakdown.
Verification. Generation is cheap. Reading, understanding, tests, and fixes are not. If you did not write the code yourself, a mental model of the system did not “grow” with it: you buy typing speed at the price of more expensive verification.
Integration. Your own code often arrives already inside your picture of the system. A model suggestion can be locally correct and globally expensive: a different error style, a different abstraction layer, a bypass of an existing API.
Debugging rabbit holes. The chain is familiar: generate → subtle bug → fix → new bug → more checking → rollback. Trust in plausible text strengthens the trap. A related story is review after an agent patch: a human must check what looks finished.
METR 2026: measurement got fragile, conclusions more cautious
In February 2026 METR explained why continuing the same experiment design became hard (We are Changing our Developer Productivity Experiment Design). Developers increasingly refuse to take part if AI is banned. Selection bias appears: the sample has fewer people who most believe in the assistant’s benefit, and fewer tasks people refuse to do “by hand.” Some participants run several agents at once, and accounting for time spent becomes unreliable. Lower participation pay may also have strengthened selection.
Raw estimates from the new run already hint at speedups in some cohorts, but the organization writes plainly: the signal is weak, and the true effect among developers and tasks “selected out” may be higher. The reader takeaway matters more than the number: early-2025 results cannot be mechanically transferred to early-2026 tools. Models, agents, and team habits changed. So did the chance of measuring the effect honestly with the old protocol.
Microsoft field experiments: thousands of developers
A large field cut is Microsoft Research’s The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. Three experiments at Microsoft, Accenture, and a Fortune 100 company, totaling 4,867 developers. In pooled data the researchers report about a +26% lift in completed tasks for those given a coding assistant. The effect is uneven: less experienced developers tended to use AI more intensively and saw larger gains.
Source: Microsoft Research.
Here the metric is different again: not minutes on one task in a familiar open-source repository, but volume of closed tasks in an organizational field. More closed tasks can mean both speedup and a shift toward work that is easier to “push through” with an assistant. So the contradiction table in the next section is not a science bug — it is different measurement axes.
Why the studies contradict each other
The central analytical answer: they measure different things and place people in different conditions.
| Factor | How it shapes the AI effect |
|---|---|
| Simple, short task | Often a strong speedup |
| Boilerplate and typical CRUD | The model is especially useful |
| Well-defined requirements | Easier to check and accept |
| Large legacy codebase | Gains melt; verification cost rises |
| Architectural decisions | Need a human |
| Good tests | Model errors are cheaper to catch |
| Weak tests | Hallucination cost rises |
| Developer experience | Changes both usage and payoff |
| Familiarity with the repository | Can reduce relative benefit |
| Model generation and mode (completion vs agent) | Effect drifts over time |
| Code review and integration | Part of the “saved” time comes back |
Add vendor conflicts of interest, sample size, observation length, and whether people may choose tasks. Then “55% faster” and “19% slower” stop being a mystery and become two points on a map of work modes. On agent mode and platform limits see also agentic engineering 2026 and why agents need a platform, not only typing speed.
The full cycle: from intent to merge
Here is a simple model of total time:
T_total = T_design + T_generation + T_review + T_debug + T_test + T_integration
AI hits T_generation with particular confidence. Sometimes it also helps T_design (quick option sketches). But it can also inflate T_review, T_debug, and T_integration — especially when the draft is foreign in style and system knowledge. The net can be positive, zero, or negative. Optimizing only generation is like speeding up the press while ignoring QA and assembly.
Hence a practical rule: measure any assistant pilot by time-to-merge, defects after merge, and review iteration count — not by autocomplete speed in the IDE. Token economics without rework economics misleads the same way a cheap model without full-cycle accounting does — see cheap LLM agent TCO.
Amplifier, not replacement: junior, middle, senior
The right comparison is not “human vs AI,” but human versus human with AI. Today the model more often plays roles: pair programmer, generator, review and docs helper, API researcher, agent on a workflow slice. It expands one engineer’s radius of action: learn an unfamiliar API faster, sketch a prototype, generate tests, parse docs. But rising personal throughput does not guarantee proportional productivity for the whole system — bottlenecks shift to problem framing, architecture, and acceptance.
Junior. The entry bar drops: error explanations, boilerplate, implementation options. The risk is the mirror image: accept code you do not understand. Without mentoring and tests, “speedup” becomes debt.
Middle. Here the assistant often delivers the most everyday value: generation, refactoring, tests, documentation, reconnaissance. This group often shows up in field lifts in task count.
Senior. AI is a multiplier: alternative architectures, fast prototype, routine automation, trade-off analysis. Seniors spot bad suggestions faster — and fall into the METR trap faster if they skimp on checking a “plausible” patch in a familiar system.
Field notes on how strong teams adopt tools without hype theater: beyond the hype. On measuring an agent in testing — 11 weeks with an AI tester.
How to run your own team A/B test
Take 20–50 real backlog tasks (not a tutorial HTTP server). Assign conditions randomly or in rotation: without AI and with AI (lock which tools and versions). Measure time to done, iteration count, errors, post-review edits, test coverage, defects after merge. Do not compare lines of code alone.
Minimum metric set:
- speed: time-to-first-working-code, time-to-completion, time-to-merge;
- quality: share of tests passed, defect rate, incidents, review rejections;
- maintainability: complexity, duplication, fit with architecture;
- economics: cost per closed task, hours per feature;
- human: cognitive load, confidence, learning effect (does the author understand their own patch a week later).
Without that contour, any external percentage is someone else’s anecdote in a pretty wrapper.
What we can already claim carefully
- AI can strongly accelerate some programming tasks — especially short and formalized ones.
- The effect depends on task type, tests, familiarity with the codebase, and tool mode.
- Faster generation does not guarantee faster overall development.
- Quality may improve under controlled conditions, but that does not automatically transfer to production.
- Experienced authors in mature repositories sometimes get less benefit than they expect — including slowdowns.
- Newer model and agent generations make 2022–2025 snapshots historical, not eternal.
- Subjective speedup often diverges from the clock.
- The correct comparison is human without AI versus human with AI, not human versus model.
FAQ
Will AI replace programmers?
Short answer: it changes the mix of work; it does not remove the need for people who set the problem, hold the architecture, and own the consequences. Producing a code draft gets cheaper; framing, verification, and system understanding get more expensive — and more valuable.
Why +55% in one study and −19% in another?
Different tasks, different metrics, different codebase context, and different moments in time. A formalized HTTP server and an issue in a million-line codebase are different sports.
Should seniors be banned from AI after METR?
No. Ban blind acceptance of patches and confusing a survey with a stopwatch. Seniors often win on prototypes and routine and lose when they skimp on checking in a familiar system.
How do you measure the effect in two weeks?
Take a batch of real tasks, lock the tools, count time to merge and defects. Two weeks is little for “forever” statistics, but enough to catch the productivity illusion.
Do agents already “fix” the METR slowdown?
METR in 2026 allows that speedups may have grown, but honestly admits: the old experiment design no longer yields a reliable estimate because of participant and task selection. Trust a pilot at home over one headline.
What should a team speed up right now?
Usually readiness criteria, tests, and review — not autocomplete speed. Otherwise you speed the press and stall the assembly line.
Further reading
Nearby on this site: AI and a 15-year project, code review in the AI era, agentic engineering 2026, agents need a platform, how teams adopt AI without the hype, 11 weeks with an AI tester, chemistry of code as a decision frame.
Conclusion
The question “who writes code faster — human or AI?” is poorly posed. The model writes a draft quickly. A human with a model may close a task faster or slower — depending on what you count as a “closed task.” Microsoft, GitHub, and METR do not contradict each other if you read them as a map of regimes: tutorial bench, lab quality, mature open source, corporate field, tool-generation change.
The final frame: AI increases a human’s throughput, turning decision-making capacity into more implemented action. But the cheaper code production becomes, the more valuable problem framing, architecture, result checking, and the ability to understand the system. Do not accelerate typing — accelerate the path from intent to an accepted change.



Comments