Skip to main content

AI can make engineering dashboards look better before business value improves. The warning signs show up in acceptance rates, work complexity, code durability, and cost.

Key Takeaways

  • Larridin’s 2026 developer productivity benchmarks say AI suggestion acceptance rates above 45% warrant investigation. The question is whether accepted code survives review, reaches production, and avoids disproportionate rework.
  • PR throughput can rise because AI is producing more low-complexity work. Complexity-adjusted velocity shows whether the team is solving harder problems faster.
  • Code turnover measured at 30 and 90 days after merge helps reveal quality problems that don’t show up when the code is first accepted or deployed.

The Illusion of Productivity

AI coding tools can produce real gains. They can also inflate the signals engineering leaders are used to watching: suggestions accepted, pull requests opened, lines changed, and stories closed.

When those numbers rise, the dashboard says productivity is improving. But the business may be getting more low-complexity code, review work, or code that needs to be rewritten later. That’s velocity theater: activity looks better while durable output and business value barely move.

A metric that improves immediately can still hide a cost that appears after the reporting period, when teams are already planning the next phase of investment.

Our guide to AI developer productivity metrics explains why traditional benchmarks need more context. The next step is spotting the patterns that make a weak gain look stronger than it is.

Where Velocity Theater Shows Up

Acceptance Rates Look Too Good

A high suggestion acceptance rate can sound like proof that an AI coding tool is working. But acceptance only shows that an engineer took the suggestion. It doesn’t show whether the code survived review, reached production, or stayed in the codebase.

According to our Developer Productivity Benchmarks 2026, acceptance rates above 45% may indicate uncritical acceptance rather than exceptional tool quality. Teams above that threshold should audit what happens after acceptance.

DX makes the broader measurement problem clear: accepted code may be heavily modified or deleted before commit. That makes acceptance rate a behavior signal, not a value metric.

PR Volume Rises but Complexity Doesn’t

Raw PR counts treat a simple configuration change and a difficult architectural change as equal units of output. AI is especially effective at producing routine, well-defined code, so PR volume can climb without a similar increase in high-value work.

Complexity-adjusted velocity corrects for that distortion. It asks whether the team is completing harder work faster, not whether AI helped create more small changes.

Quality Problems Appear Later

Some AI-generated code passes review and tests but still creates maintenance problems. The issue may not become visible until weeks or months after the code was written.

Code turnover helps catch that delayed cost. It measures how much merged code is reverted, deleted, or substantially rewritten within 30 or 90 days. A team can look healthy at deployment and still have a turnover problem that becomes clear after the dashboard has already counted the work as a win.

Tool Costs Rise Faster Than Durable Output

Engineering AI spend includes licenses, token usage, implementation, training, review time, and rework. If usage and code volume rise while complexity-adjusted output and durability stay flat, the organization is paying more without creating a stronger engineering system.

That’s why utilization can’t carry the ROI case. The value has to show up in durable, quality-adjusted output.

A 5-Dimension Test for Real Engineering Value

Larridin’s engineering AI productivity measurement framework asks five questions:

  • Adoption: Are the right engineers using the tools consistently for tasks where the tools can help?
  • AI code share: How much committed code is AI-assisted, and where is it concentrated?
  • Complexity-adjusted velocity: Is the team completing more difficult, valuable work?
  • Code quality: Does AI-assisted code survive review and stay in the codebase without increasing defects, turnover, or incidents?
  • ROI: Does the measurable value exceed the full cost of the tools and the rework they create?

A productivity claim becomes credible when adoption and code share rise alongside complexity-adjusted output, stable quality, and positive ROI.

What Engineering Leaders Should Report

Boards and finance leaders don’t need another tool-usage dashboard. They need a defensible explanation of what changed after the investment.

In our work with enterprises, we’re seeing boards ask for clearer evidence that AI spending is producing P&L-level value. Engineering leaders should be ready to report:

  • The pre-AI baseline and the measurement period
  • Adoption and AI code share
  • Complexity-adjusted output
  • 30- and 90-day quality and turnover results
  • Net ROI after tool, implementation, and rework costs

That tells a stronger story than “engineers are shipping faster.” It shows whether the speed created durable business value.

Frequently Asked Questions

How do we distinguish real AI productivity gains from velocity theater?

Real gains combine higher complexity-adjusted output with stable or improving quality and positive ROI. Velocity theater happens when acceptance, PR volume, or code share rises without a similar improvement in durable output.

Is an AI code acceptance rate above 45% always bad?

No. It’s a reason to investigate, not an automatic failure. Check whether accepted suggestions survive review, reach production, and avoid disproportionate edits, defects, or turnover.

How long does it take for AI code quality problems to surface?

There’s no universal timeline. Some issues appear during review or testing, while others take weeks or months to become visible. Tracking code turnover at 30 and 90 days gives leaders a practical way to catch delayed rework.

What should engineering leaders tell boards about AI productivity?

Report changes in complexity-adjusted output, code quality, durability, and ROI against a clear pre-AI baseline. Adoption and usage belong in the picture, but they shouldn’t be presented as business value.

Measure Value, Not Velocity Theater

Larridin measures adoption, AI code share, complexity-adjusted velocity, code quality, and ROI so engineering leaders can see whether AI is creating durable value or simply creating more code.

Book a discovery call to see what your engineering AI investment is actually producing.