Microsoft’s 2026 field study gives enterprise leaders stronger evidence that agentic coding tools can increase engineering output. It also makes the limit clear: more merged pull requests aren’t the same as more business value, and the added output has to justify the cost.
AI coding productivity research has produced mixed results because studies measure different tools, tasks, developers, and outcomes. Microsoft’s study addresses one important gap: what happens when agentic command-line coding tools are used in a large software organization.
The researchers studied tens of thousands of engineers during Microsoft’s early-2026 rollout of Claude Code and GitHub Copilot CLI. Instead of asking developers whether they felt faster, the study used merged pull requests as a concrete measure of output.
That makes the study useful for enterprise leaders, but not universal. It measures one organization, two tools, and one output metric. The useful question is what the result tells leaders to measure in their own environment.
The study estimated that adopters merged 24% more pull requests per engineer per day than they would have without the tools. The 95% confidence interval ranged from 14.5% to 33.7%.
The lift also held across the roughly four-month observation window. The researchers didn’t find a statistically meaningful decline between February and March-April, suggesting the result wasn’t simply a short-lived novelty effect.
That’s meaningful evidence that agentic coding tools can move a concrete engineering output metric at scale.
But it is still one metric.
A merged pull request is not the same as a feature shipped, a customer problem solved, or durable code in production. The paper makes that distinction itself.
It doesn’t measure whether the additional PRs improved code quality, reduced incidents, lowered rework, or created more business value. In its conclusion, the paper calls quality the next open question.
That matters because higher throughput can look positive while downstream work gets worse. Larridin’s Developer Productivity Benchmarks 2026 recommend reading velocity alongside code quality, AI code share, adoption, and cost rather than using one volume metric as the productivity score.
The 24% figure also shouldn’t become a universal benchmark for AI coding ROI. It reflects Microsoft’s environment, rollout, tools, and measurement approach. Your own baseline is still a better reference point for determining whether AI changed output.
The paper opens with a straightforward warning: token spend can run into millions of dollars annually at organizational scale.
It also cites an extreme example from Meta, reported by Fortune, in which one employee’s usage could have cost more than $1.4 million in a month. That wasn’t a Microsoft engineer or a typical user in the study. The example simply shows how large consumption can become at the far end of the distribution.
For enterprise leaders, the important question is whether added output is worth what it costs.
A team that merges more PRs while token spend rises sharply may still have a strong return. Another team can show the same throughput gain while producing more rework, defects, or code turnover. The output number alone cannot distinguish the two.
Start with three views.
Track the engineering result AI is supposed to improve: merged PRs, cycle time, complexity-adjusted throughput, or another delivery measure that fits the team.
Check whether the additional output holds up. Code turnover, rework, change failure rate, incidents, and review patterns can show whether faster output is creating downstream cost.
Track AI spending by tool, team, engineer, agent, or use case rather than relying on a total invoice. Larridin’s Token Spend and Insights platform attributes AI spending across tools and workflows, while the AI Dev Productivity platform connects AI activity with delivery and quality signals.
The goal is to see whether the added output remains useful after cost and quality are included.
No. The study found a 24% lift in merged PRs in Microsoft’s environment. It didn’t find a 24% financial return or business-value increase. Use the result as evidence that agentic coding tools can affect throughput, then measure your own baseline, cost, and downstream outcomes.
Not within the study window. The researchers found no statistically meaningful decline between the later periods they compared. That supports the conclusion that the lift persisted across the roughly four-month observation period.
No. The Microsoft paper cites a Fortune report about an extreme Meta user to illustrate the scale token consumption can reach. It’s not a Microsoft employee or a representative usage level.
The paper says its observation window ended April 29, 2026, and that an internal announcement shortly afterward directed most affected engineers to move from Claude Code to Copilot CLI. The study doesn’t say the decision was caused by cost or by the productivity findings, so those events shouldn’t be treated as a cause-and-effect story.
Establish a baseline before or early in the rollout, then track delivery, quality, and AI spend over the same period. Larridin’s AI Dev Productivity and Token Spend & Insights platforms connect those signals so leaders can see what changed and what the change cost.
The Microsoft study shows that agentic coding tools can produce a measurable increase in merged PRs. The next step is determining whether that extra output is durable and worth the cost.
Larridin connects AI usage, engineering delivery, quality, and spend so leaders can evaluate the full result rather than stopping at a throughput number.
Book a discovery call to connect your AI productivity and cost signals.