Larridin Blog

AI vs. Human Developer Productivity: How to Compare Output

Written by Larridin | Aug 7, 2026

METR found a slowdown among experienced open-source developers using early-2025 tools. A Google trial found a speedup on an enterprise task. The contradiction shows that broad studies can inform your measurement plan, but they can’t tell you what AI is doing in your environment.

Key Takeaways

  • Research on AI developer productivity produces different results because tools, tasks, developers, and measurement methods differ.
  • Developer self-reports provide useful experience data, but they can’t replace system-measured delivery, quality, durability, and cost signals.
  • A trustworthy comparison uses matched work, AI attribution, complexity-adjusted throughput, and downstream outcomes for AI-assisted and human-only work.

Why Research Doesn’t Produce One Universal Answer

The METR randomized controlled trial involved 16 experienced open-source developers completing 246 real tasks in repositories they knew well. With early-2025 AI tools available, they took 19% longer. Before the study, they expected AI to make them 24% faster. Afterward, they estimated that it had made them 20% faster.

That gap shows why perceived speed and measured completion time can diverge. But it doesn’t prove that AI always slows experienced developers.

METR later said newer tools would likely produce quicker results than the original study. Its follow-up experiment showed some evidence of faster work, but selection effects and difficulty measuring parallel agent use made the size of the improvement unreliable.

A separate randomized trial with 96 Google engineers found that AI shortened time on a complex enterprise-grade task by about 21%. The confidence interval was wide, and the researchers cautioned against assuming the result would transfer to other tools, tasks, or environments.

These studies show that the result depends on the work, the developer, the tooling, and how productivity is measured.

The 5 Requirements for a Trustworthy Comparison

1. Start With Comparable Work

Compare the same team before and after AI adoption, or compare similar teams completing similar work during the same period.

Control for factors that can distort the result, including repository maturity, task type, developer experience, release pressure, and major process changes. A team handling routine maintenance shouldn’t be compared directly with one rebuilding a core architecture.

Use industry benchmarks for context, not as the primary baseline. The strongest comparison is usually the team’s own historical performance on comparable work.

2. Attribute Work to AI-Assisted and Human-Only Sources

You can’t compare AI-assisted and human-only output unless you know where AI contributed.

Track AI code share at the line, commit, or pull request level. Then apply the same delivery and quality measures to AI-assisted and human-only cohorts.

Attribution needs to follow committed work, not suggestion volume. An accepted completion that never reaches the main branch doesn’t count as delivered output.

Larridin’s Developer Productivity platform connects AI activity with repositories, workflows, delivery, quality, and cost signals so leaders can compare the work using one data model.

3. Adjust Throughput for Complexity

Raw pull request, commit, and line counts favor high-volume work. AI can generate boilerplate, configuration changes, tests, and routine fixes quickly, but that doesn’t make each unit equally valuable.

Complexity-adjusted throughput weights output by the difficulty of the work. This helps leaders see whether AI is increasing meaningful delivery or mainly expanding the volume of simple changes.

Compare AI-assisted and human-only work within similar complexity bands. That prevents a team using AI for routine tasks from appearing more productive than a team handling fewer, harder changes.

4. Pair Speed With Durability and Quality

Completion time matters only when the output makes it through review and production.

Track code turnover, reverts, review burden, defects, incidents, security findings, and remediation for both cohorts. Use consistent 30- and 90-day windows so recent code has the same opportunity to fail or be rewritten.

Larridin has reported sharply different results across customer environments. In one, AI-assisted code had a 0.2% revert rate compared with 15.16% for human-only code. In another, 30-day turnover rose 64.7% alongside increasing AI code share. Those results shouldn’t be generalized, but they show why every organization needs its own comparison.

Our revert rate measurement framework explains how to compare durability across AI-assisted and human-only code.

5. Connect the Comparison to Delivery and Cost

It doesn’t matter whether developers generated more code. What’s important is whether the organization delivered more useful work at an acceptable cost and risk level.

Compare lead time, deployment frequency, change fail rate, deployment rework rate, and recovery time before and after AI adoption. Read those signals beside total AI coding spend, review labor, remediation, and durable output.

A team is gaining ground when AI-assisted work helps key changes reach production faster while quality stays stable or improves. Rising activity without better delivery, durability, or cost efficiency isn’t a productivity gain.

Use Self-Reports as Context, Not the Score

Developer surveys still matter. They can identify friction, satisfaction, confidence, workflow fit, and training needs that system data won’t fully explain.

Use them to understand why measured outcomes differ across teams. A group with low adoption and strong delivery may have a tool-fit problem. A team with high usage and weak durability may need better review practices or AI proficiency.

But don’t ask developers to convert their experience into an ROI estimate. Pair qualitative feedback with system-measured outcomes.

Frequently Asked Questions

Is AI-assisted development more productive than human-only development?

It depends on the task, tool, developer, workflow, and outcome being measured. Controlled studies have found both slowdowns and speedups. The reliable answer for an organization comes from comparing matched work in its own environment.

How do we reduce perception bias in AI productivity comparisons?

Use system data for the primary comparison. Track AI attribution, complexity-adjusted throughput, delivery time, rework, durability, incidents, and cost. Use developer feedback to explain the results rather than replace them.

What is the best baseline for comparing AI-assisted and human-only work?

Use the same team’s performance on comparable work before AI adoption when possible. A matched cohort of similar teams or repositories can also work. Avoid comparing one team directly with a broad industry average.

Should we track AI-assisted and human-only productivity continuously?

Yes. Tools, models, workflows, and team proficiency change quickly. A quarterly or annual snapshot can miss a shift in quality, cost, or delivery performance as AI code share grows.

Measure What AI Changes in Your Engineering System

Larridin connects AI attribution, complexity-adjusted throughput, code durability, delivery performance, and cost. Leaders can compare AI-assisted and human-only work using measured outcomes instead of relying on activity counts or perceived speed.

Book a discovery call to build your AI developer productivity comparison.