Does AI coding actually work? The honest answer is three honest answers, because different kinds of evidence measure different things. A controlled trial measures what happened to specific developers. An industry survey measures what organizations associate with adoption. An operator report measures what one company says about itself. Read across all three and a usable picture emerges — narrower than the hype, larger than the dismissal.
What the controlled study found
METR's July 2025 randomized trial assigned 246 real issues across 16 experienced open-source developers to AI-allowed or AI-disallowed work. With early-2025 tools, the AI-allowed tasks took 19 percent longer — and the developers believed they had been about 20 percent faster. Perception and measurement pointed in opposite directions.
METR's February 2026 update complicates the clean story in both directions. The follow-on estimates an 18 percent speedup for returning developers and 4 percent for new ones — but participants had begun refusing AI-free assignments, so METR itself calls the signal weak and is redesigning the study. The fairest reading: the early-2025 slowdown was real in its setting, the effect has likely moved since, and nobody has a clean current measurement.
What the industry survey found
DORA's 2025 report describes AI as an amplifier: adoption is broad, AI use correlates with modest gains in throughput, and the same organizations report low trust in generated work alongside continued instability. The survey measures associations across thousands of respondents — direction, not controlled effects — and DORA maintains a public errata record for it.
What operators report at scale
The strongest claims come from organizations reporting on themselves. Stripe's president reported roughly 7,000 pull requests from Minions in one week — about 30 percent of that week's PRs. Uber's factory account reports token economics and adoption across layers. Spotify's account says coding is no longer the constraint — and that review, standardization, and context became the new ones. Ramp's SWE-Bench rebuilt the benchmark from production tasks because public ones did not match its work.
These are first-party reports: real evidence with a known bias toward the flattering number. They are also, at present, the only org-level scale data that exists.
How to read the contradiction
The disagreement is smaller than it looks because the denominators differ:
- Benchmarks (SWE-bench Lite, Terminal-Bench 2.0, HarnessTax) measure capability on bounded tasks — what an agent can do, not what a team gets.
- The METR trial measured individual developers on repositories they knew intimately, with tools from early 2025 — before current harnesses and models.
- Operator reports measure merged-work share inside prepared systems: task briefs, scoped context, checks, and human review. The system, not the model alone, produces the number.
A working rule: trust labeled evidence over averages, and treat "does it work" as a question about a specific task class inside a specific system. The figures that will actually decide your next quarter are the ones you measure — the metrics guide shows how to instrument a bounded pilot so your evidence, not the discourse, answers the question.