Skip to main content

Engineers can report the same productivity gains six months later even as their developer experience gets worse. A longitudinal study shows why a second measurement can matter as much as the first.

Key Takeaways

  • In a six-month study, the share of engineers reporting improved productivity stayed at 84%, while reports of worsened developer experience increased.
  • Engineers also reported spending less time writing code and more time reviewing, testing, and directing AI-generated work.
  • A one-time productivity check can miss both changes. AI coding measurement needs to show whether early results hold as usage and engineering work evolve.

The Headline Productivity Number Didn’t Change

Researchers Annie Vella and Kelly Blincoe at the University of Auckland surveyed professional software engineers twice, six months apart, in October 2024 and April 2025. 95 engineers completed both surveys and formed the matched group used for the longitudinal analysis.

The headline productivity result was remarkably stable: 84% said AI coding tools improved their productivity at both check-ins.

If the organization had asked only that question, either survey would have produced essentially the same story. But other parts of the experience were moving.

Among matched respondents, the share reporting that AI tools had worsened their developer experience in at least one dimension increased from 14% to 27% over the six months. Flow state and cognitive load worsened, while feedback loops improved.

That’s the productivity-experience paradox the researchers identified. The measurement lesson is that productivity and developer experience can diverge over time even when the headline productivity result barely moves.

The Work Was Changing Too

The second survey captured another shift that a simple productivity question wouldn’t show.

By the later check-in, 82% of participants reported spending less time writing code. Their effort had shifted toward reviewing, testing, checking, and directing AI-assisted work.

As engineers described that change, the researchers identified a category they call supervisory engineering work: directing AI tools, evaluating what they produce, and correcting or integrating the output.

That includes work such as:

  • Giving an AI tool direction and additional context
  • Reviewing generated code
  • Testing whether the output works
  • Deciding what to accept, revise, or discard
  • Correcting AI-generated work before it ships

The term comes from this study and isn’t a standardized industry metric, but the underlying shift matters whether or not organizations adopt the label.

If engineers are writing less code themselves and spending more time supervising AI-generated work, a measurement system based on older assumptions about where engineering effort goes can miss that change.

Why One-Time Measurement Falls Short

This is where this study adds something a snapshot cannot.

A point-in-time survey can tell you how engineers perceive an AI coding tool today. It can’t tell you whether that perception holds as people become more experienced with the tool, use it for different work, or take on more review and verification.

The first survey also can’t show whether the nature of the work is gradually changing.

That makes repeated measurement useful for questions such as:

  • Is the productivity benefit holding up?
  • Is developer experience improving or deteriorating?
  • Are engineers spending less time creating and more time verifying?
  • Has AI shifted effort into work that the existing dashboard does not capture well?

The goal is to avoid treating the first positive reading after rollout as a permanent answer.

What to Revisit as AI Coding Use Matures

Engineering leaders don’t need a new metric for every behavior identified in the study. They do need a way to tell whether the picture is changing.

A practical follow-up can combine two kinds of evidence.

  1. Operational engineering data can show how AI-assisted work affects delivery, quality, rework, verification, and cost over time.
  2. Periodic developer feedback can show whether cognitive load, flow, and the experience of reviewing and directing AI are changing as usage matures.

Larridin’s Agent Effectiveness addresses the operational side by tracking agent-assisted engineering behavior and outcomes over time. Developer experience still needs to be measured separately through periodic feedback.

That distinction matters. The Auckland study shows that perceived productivity and developer experience can move differently over time. That’s why AI coding measurement needs a return visit. Recheck both the engineering outcomes and the developer experience to see whether the original results still hold.

Frequently Asked Questions

What’s the productivity-experience paradox?

The researchers use the term to describe a pattern in which perceived productivity from AI coding tools stayed high while negative effects on parts of developer experience became more common over six months.

Why does AI coding productivity need to be measured more than once?

A single measurement cannot show whether early gains persist or whether developer experience and engineering work change as teams use AI coding tools longer.

What’s supervisory engineering work?

It’s the researchers’ term for work involved in directing AI tools, evaluating their output, and correcting or integrating what they produce. It came from this study and isn’t a standardized industry metric.

Does the study prove AI coding tools hurt developer experience over time?

No. The study used self-reported data from a voluntary sample, and the findings shouldn’t be treated as a universal benchmark. It identifies a longitudinal pattern worth checking for in other engineering organizations.

See How AI-Assisted Engineering Changes Over Time

Larridin connects AI-assisted engineering activity with delivery, quality, verification, rework, and spend so teams can track whether AI coding performance continues to improve after rollout.

Talk to an expert.