Skip to main content

Microsoft’s 2026 field study gives enterprise leaders stronger evidence that agentic coding tools can increase engineering output. It also makes the limit clear: more merged pull requests aren’t the same as more business value, and the added output has to justify the cost.

Key Takeaways

  • Microsoft researchers used direct telemetry from tens of thousands of engineers during the company’s early-2026 rollout of Claude Code and GitHub Copilot CLI. The study gives leaders real-world evidence about agentic coding tools rather than relying only on surveys or lab tasks.
  • The headline result is a 24% lift in merged pull requests, not a 24% ROI gain. The paper treats merged PRs as an output measure and explicitly says they aren’t the same as the value delivered.
  • Cost belongs in the same ROI analysis. The paper warns that token spend can reach millions of dollars annually at organizational scale, so higher output needs to be read alongside spend, quality, and rework.

Why This Study Matters

AI coding productivity research has produced mixed results because studies measure different tools, tasks, developers, and outcomes. Microsoft’s study addresses one important gap: what happens when agentic command-line coding tools are used in a large software organization.

The researchers studied tens of thousands of engineers during Microsoft’s early-2026 rollout of Claude Code and GitHub Copilot CLI. Instead of asking developers whether they felt faster, the study used merged pull requests as a concrete measure of output.

That makes the study useful for enterprise leaders, but not universal. It measures one organization, two tools, and one output metric. The useful question is what the result tells leaders to measure in their own environment.

What the Study Found

The study estimated that adopters merged 24% more pull requests per engineer per day than they would have without the tools. The 95% confidence interval ranged from 14.5% to 33.7%.

The lift also held across the roughly four-month observation window. The researchers didn’t find a statistically meaningful decline between February and March-April, suggesting the result wasn’t simply a short-lived novelty effect.

That’s meaningful evidence that agentic coding tools can move a concrete engineering output metric at scale.

But it is still one metric.

What the 24% Lift Does Not Prove

A merged pull request is not the same as a feature shipped, a customer problem solved, or durable code in production. The paper makes that distinction itself.

It doesn’t measure whether the additional PRs improved code quality, reduced incidents, lowered rework, or created more business value. In its conclusion, the paper calls quality the next open question.

That matters because higher throughput can look positive while downstream work gets worse. Larridin’s Developer Productivity Benchmarks 2026 recommend reading velocity alongside code quality, AI code share, adoption, and cost rather than using one volume metric as the productivity score.

The 24% figure also shouldn’t become a universal benchmark for AI coding ROI. It reflects Microsoft’s environment, rollout, tools, and measurement approach. Your own baseline is still a better reference point for determining whether AI changed output.

Cost Is Part of the Same ROI Question

The paper opens with a straightforward warning: token spend can run into millions of dollars annually at organizational scale.

It also cites an extreme example from Meta, reported by Fortune, in which one employee’s usage could have cost more than $1.4 million in a month. That wasn’t a Microsoft engineer or a typical user in the study. The example simply shows how large consumption can become at the far end of the distribution.

For enterprise leaders, the important question is whether added output is worth what it costs.

A team that merges more PRs while token spend rises sharply may still have a strong return. Another team can show the same throughput gain while producing more rework, defects, or code turnover. The output number alone cannot distinguish the two.

What to Measure Alongside PR Lift

Start with three views.

1. Output

Track the engineering result AI is supposed to improve: merged PRs, cycle time, complexity-adjusted throughput, or another delivery measure that fits the team.

2. Quality

Check whether the additional output holds up. Code turnover, rework, change failure rate, incidents, and review patterns can show whether faster output is creating downstream cost.

3. Cost

Track AI spending by tool, team, engineer, agent, or use case rather than relying on a total invoice. Larridin’s Token Spend and Insights platform attributes AI spending across tools and workflows, while the AI Dev Productivity platform connects AI activity with delivery and quality signals.

The goal is to see whether the added output remains useful after cost and quality are included.

Frequently Asked Questions

Should we use the 24% PR lift as our expected ROI from AI coding agents?

No. The study found a 24% lift in merged PRs in Microsoft’s environment. It didn’t find a 24% financial return or business-value increase. Use the result as evidence that agentic coding tools can affect throughput, then measure your own baseline, cost, and downstream outcomes.

Did the productivity lift fade after the initial rollout?

Not within the study window. The researchers found no statistically meaningful decline between the later periods they compared. That supports the conclusion that the lift persisted across the roughly four-month observation period.

Was the $1.4 million per month example a Microsoft engineer?

No. The Microsoft paper cites a Fortune report about an extreme Meta user to illustrate the scale token consumption can reach. It’s not a Microsoft employee or a representative usage level.

Why did Microsoft discontinue most Claude Code licenses?

The paper says its observation window ended April 29, 2026, and that an internal announcement shortly afterward directed most affected engineers to move from Claude Code to Copilot CLI. The study doesn’t say the decision was caused by cost or by the productivity findings, so those events shouldn’t be treated as a cause-and-effect story.

How do we measure productivity and cost together?

Establish a baseline before or early in the rollout, then track delivery, quality, and AI spend over the same period. Larridin’s AI Dev Productivity and Token Spend & Insights platforms connect those signals so leaders can see what changed and what the change cost.

Connect Output to Cost and Quality

The Microsoft study shows that agentic coding tools can produce a measurable increase in merged PRs. The next step is determining whether that extra output is durable and worth the cost.

Larridin connects AI usage, engineering delivery, quality, and spend so leaders can evaluate the full result rather than stopping at a throughput number.

Book a discovery call to connect your AI productivity and cost signals.