Skip to content
AI Measurement Guide / Engineering AI ROI
The Larridin ROI Measurement Framework

How to measure
and prove AI ROI
in Engineering.

Agree on what AI should improve for the business. Work backward to the evidence, then measure progress and adjust as you go.

Guide + interactive worksheet
THE LARRIDIN ROI MEASUREMENT FRAMEWORK

Start with what matters
to the business.

Align on the outcome. Work backward to the evidence that would demonstrate it. Use telemetry to check progress, understand what changed, and decide what to adjust.

  1. Align on the goal

    Engineering, Product, and Finance agree on the business objective, how Engineering can advance it, and what success looks like.

    Align goals and contribution
  2. Work backward to the evidence

    Identify the work and milestones that should advance the goal. Set the baseline, costs, and quality requirements you will compare against.

    Connect work to outcomes
  3. Measure, learn, and adjust

    Use telemetry to follow delivery, cost, and quality. Validate the business benefit with your partners and use the findings to improve the next investment.

    Turn evidence into a decision

Review the evidence together. Adjust the workflow, investment, or goal as you learn.

Your engineers are using AI. Some work is moving faster, and the AI bill is growing. The question is whether that investment is advancing the priorities your business cares about.

Start by agreeing on the business objective and Engineering’s role. It might be delivering a customer commitment sooner, completing a migration within budget, or reducing the cost of recurring failures.

Then work backward: what needs to change in engineering, what evidence would show that change, and how will you know the business benefited? Telemetry makes progress visible along the way, so you can test your assumptions and adjust while the work is still underway.

Start with business objectives.
Define Engineering’s role.

Agree with Product and Finance on the business objective first. Then identify how Engineering can advance it: the customer capability to deliver, the migration to complete, or the operational problem to resolve. That connection determines where AI could help and what evidence you need.

Make the relationship explicit: the business objective, Engineering’s contribution, and the evidence of progress. Agree on ownership, timing, and what success looks like before selecting metrics or evaluating a rollout.

Choose the engineering improvement that supports the objective.

↗

More capacity

More comparable, accepted work with stable staffing and quality.

◷

Earlier delivery

A priority initiative reaches its milestone sooner because its critical path shortened.

−$

Lower expenditure

A verified reduction in contractors, overtime, or another expense.

✓

Better reliability

Fewer failures and a measurable reduction in their operational or customer cost.

Count each benefit once. If recovered capacity both avoids contractor work and accelerates delivery, establish whether those are separate benefits before counting both.
Larridin starting points

Once the business objective and Engineering’s contribution are agreed, choose the measures that will show progress within a consistent scope and period.

Measure substantive work

AI makes it easy to produce more code. Engineering leaders need to know whether the organization is completing more meaningful work.

PR counts help describe activity, but the units vary. Ten small configuration changes and one architectural change represent different work. Changing PR-splitting practices can move the count without changing delivery.

Larridin’s Engineering Output provides a complementary measure. It scores merged changes for complexity, includes a bounded adjustment for size, and reduces the score for missing tests and AI slop. Teams can inspect the underlying PRs to understand what moved the aggregate. [1]

Use the trend with consistent repositories, teams, and completed weeks. Check whether changes reflect more completed work, a different complexity mix, or improved quality at merge.

Two boundaries matter. First, merged work may not have reached customers. Second, implementation complexity is not customer value. A small change can unlock a major customer; a difficult feature can fail to find users.

Keep the output measure alongside releases and initiative milestones. Together, they answer how much work was completed and what that work accomplished.

Larridin starting point

Inspect underlying PRs and compare consistent repositories, teams, and completed weeks. Pair the trend with release and initiative evidence.

Show where the additional capacity went?

Suppose output rises 40%. The next question is where that increase landed.

Larridin’s Innovation Rate classifies work as Features, Bug Fixes, or KTLO—keeping the lights on—and defaults to weighting the mix by Engineering Output. With aligned scope and classification coverage, multiplying total Output by Feature share gives a useful derived view of feature output. [2]

Work allocation

Where additional output goes

Consider a team whose total Output rises from 100 to 140 while Feature Output falls from 60 to 56.

Feature OutputBug Fixes + KTLO
Baseline
6040
100
Current
5684
140
+40%Total Engineering Output
−4 pointsFeature Output · 60 → 56

Feature share moved from 60% to 40%. The business interpretation depends on the work: maintenance may retire a fragile service or repair poor changes.

The team completed more work, but less feature work. That is an allocation finding, not automatically a bad result.

The additional maintenance may have retired a fragile service, addressed security issues, or removed a constraint on future delivery. It may also represent repairs caused by poor changes. The business interpretation depends on the work.

Name the intended allocation before celebrating the gain. If the objective was faster product delivery, show progress against that objective. If the objective was debt reduction, show the systems retired, recurring problems removed, or operating burden reduced.

Larridin’s leadership playbooks connect this discussion to initiative completion, due dates, planned and unplanned work, and PR-to-work-item linkage. Those connections make it possible to follow output into a specific commitment. Unlinked work is a gap in that explanation. [3]

Larridin starting points

Read the output trend alongside work allocation and the commitments that work advanced.

Check whether the whole delivery system improved

AI may shorten implementation while creating more work for reviewers, testers, or release teams.

Track both PR cycle time and the path into production. Larridin’s Velocity view measures first commit to merge; its CI/CD views add deployment frequency, change lead-time stages, change failure rate, and restoration time. These boundaries are different and should remain explicit. [4]

A team can merge faster without releasing sooner. Equally, faster releases can still fail to advance a launch if another dependency remains on the critical path.

Imagine a team that cuts implementation from four days to two. Review waits grow by two days, and the release date stays unchanged. The team has improved one stage, but has not demonstrated an end-to-end delivery gain.

The bottleneck moved.

Suppose implementation falls from four days to two, but review takes two days longer.

Before
Implementation 4d
Review 1d
Release 1d
After AI
Implementation 2d
Review 3d
Release 1d
6 days → 6 daysNo end-to-end delivery gain demonstrated.

This is useful evidence. It tells leadership where to invest next: review capacity, clearer requirements, testing, or deployment automation.

WorkGraph can add context by showing how observed engineering attention shifts among development, design, quality, incidents, operations, collaboration, and administration. Its shares describe observed attention, not contracted hours, so they should not be converted directly into payroll savings. [5]

Larridin starting points

Keep first-commit-to-merge and production delivery boundaries explicit. Follow the critical path into the actual release milestone.

Make quality part of the claim

Additional output creates value only if the work is fit for its purpose. Verification, maintenance, and production failures can consume an apparent gain.

Use two kinds of quality evidence.

Signals available around merge identify risk: expected tests, code-quality patterns, review engagement, and change size. Later evidence shows what happened: rework, reversions, failed deployments, and incidents.

Larridin’s quality metrics include AI Slop Index, Unit Tests, review signals, and 30-day and 90-day rework. Engineering Output already incorporates missing-test and slop penalties, but those early assessments do not replace later outcome checks. [2]

Compare work at similar ages. A PR merged yesterday has had less opportunity to be rewritten than one merged three months ago. Mark immature rework windows as provisional.

Interpret the signal before assigning a cost. Rewritten code can reflect a defect, a changed requirement, or an intentional redesign. Review comments are evidence of discussion, not a complete measurement of review effort.

For production consequences, examine change failure rate, restoration time, incident severity, and response burden. Establish the relevant service and change linkage before attributing a reliability change to AI. [4] [6]

A credible productivity claim should survive these checks. If it does not, the additional verification or repair work belongs in the investment decision.

Larridin starting points

Use merge-time signals and mature downstream outcomes together. Investigate the cause before valuing rework or attributing failures.

Establish what AI contributed

AI involvement and incremental AI impact are separate questions.

Larridin distinguishes AI PR Share, AI Code Share, and AI Output Share. They measure involvement using PR counts, attributable added lines, and Output respectively. Their percentages need not match. [7]

The attribution rule matters. Under the documented method, a human-authored PR with direct AI evidence contributes its full Output to AI Output; inferred-only attribution contributes proportionally. Therefore, a 70% AI Output Share does not mean AI independently performed 70% of the work or increased productivity by 70%. [7]

Use attribution to identify the work being studied. Then estimate how outcomes changed relative to a credible baseline.

Begin with the same team doing comparable work before and after an intervention. Record staffing, repository scope, work mix, release policies, and other changes that might explain the result.

Where possible, strengthen the comparison with a similar team or a staggered rollout. If both groups improve, distinguish the improvement common to both from the additional change in the intervention group. This still requires checking that their earlier trends and operating conditions are comparable.

For bounded tasks, controlled evaluations can provide stronger evidence. For broad rollouts, triangulate output, delivery, quality, and developer feedback rather than relying on one comparison.

Match the claim to the evidence. “Output increased after the rollout” describes an observation. “The rollout contributed to an estimated increase” requires a stronger comparative argument. Reserve firm causal claims for a design that can support them.

OBSERVATION

“Output increased after the rollout.”

↓
STRONGER COMPARATIVE EVIDENCE

“The rollout contributed to an estimated increase.”

Also track collection coverage. Better instrumentation can increase measured AI involvement even when behavior has not changed.

Larridin starting points

Use attribution to select the work being studied; use comparative evidence to investigate incremental impact.

WHAT THE ORIGINAL DATA CAN TELL US

Higher spend is not proof of higher return.

Larridin’s spend/output curves vary by AI-attribution cohort. Points are spend-band medians normalized to each cohort’s lowest-spend band, not before-and-after productivity gains. The two AI-native cohorts share a company; the low-AI cohort is from another company with partial repository coverage.

These associations do not establish causation or financial ROI.

Explore the curves and uncertainty bands ↗
ORIGINAL RESEARCH · LARRIDIN DATA

What does AI coding actually cost?

Weekly billed AI-coding spend among observed engineers who both merged code and incurred billed spend:

25th percentile$86
Median$213
75th percentile$490
90th percentile$911

Four complete weeks ending August 2, 2026. Provider-invoiced spend; flat-fee consumption excluded. Median 95% bootstrap interval: $178–$262. Use as context, not a budget target.

Source: Larridin Data, Finding 1 ↗ Methodology ↗

Calculate three different economic measures

A single dollar figure can hide several different ideas. Keep delivery efficiency, capacity value, and realized financial ROI separate.

Delivery efficiency

Delivery efficiency asks what it costs to produce a comparable unit of accepted work.

Total delivery cost per unit = relevant labor, AI, and other delivery costs ÷ comparable accepted output.

An AI-spend-only Cost per Output measure helps determine whether tooling expense is growing faster than output. It does not capture the entire cost of delivery. Review labor, infrastructure, and other costs need a consistent treatment in the broader calculation. Larridin’s leadership playbooks use Cost per Output and Cost per Outcome to guide investment decisions. [8]

Capacity value

Capacity value estimates the economic equivalent of additional work the team can perform.

Estimated capacity value = incremental capacity × an explicit valuation rate.

Label the baseline and assumptions. Do not convert Output points into hours without a separately validated conversion. If using Larridin’s AI Capacity Added, preserve its definition as estimated manual-effort equivalence from observed assistant conversations. It is not a direct record of saved working hours, and overlapping work must not be added again as a separate engineering benefit. [9]

Realized financial ROI

Realized financial ROI asks what verified incremental benefit the company captured relative to the additional investment.

Realized ROI = (incremental realized benefit − total incremental investment) ÷ total incremental investment.

Keep the period consistent. Include the relevant incremental licenses, usage, infrastructure, integration, enablement, and additional human verification or remediation costs. Avoid charging the same labor twice.

Cost evidence also has different strengths. Larridin’s Agent Effectiveness documentation distinguishes provider-billed spending from usage-priced session estimates. An estimate can support a workflow comparison while still needing reconciliation for financial reporting. [10]

Larridin starting points

Reconcile the cost basis. Keep delivery efficiency, estimated capacity value, and realized financial ROI separate.

Follow the benefit into the business

Consider these scenarios. Each calls for a different claim and different evidence.

A product team creates capacity

The team produces 25% more comparable output with stable staffing and quality. Some of the increase reaches features, but customer release dates remain unchanged.

The defensible claim is increased engineering capacity and completed work. It is premature to claim launch acceleration or incremental revenue. Follow which additional commitments the team completes with that capacity.

A migration team avoids expenditure

The company has an approved, executable plan to spend $200,000 on contractors for a defined migration. After introducing AI-assisted workflows, it completes the same scope with $80,000 of contractor spending. AI tooling, enablement, and incremental internal effort total $40,000.

If Finance confirms the avoided $120,000 expense and the scope and quality are equivalent:

A migration team avoids expenditure

Consider a migration with a $200,000 contractor plan, $80,000 in actual contractor spending, and $40,000 in incremental AI-program costs.

Avoided expense$120k
Incremental investment−$40k
Net benefit$80k
REALIZED ROI200%

($120k − $40k) ÷ $40k
Only if Finance confirms the avoided expense, scope and quality are equivalent, and the delivery analysis supports AI’s contribution.

Approved contractor plan: $200k. Actual contractor spending: $80k. AI tooling, enablement, and incremental internal effort: $40k.

The canceled expenditure provides the financial evidence. The delivery analysis must still establish AI’s contribution and rule out scope reduction or uncounted effort as the explanation.

A customer initiative ships sooner

A team releases a capability four weeks earlier than a credible baseline. The delivery record shows which critical-path tasks accelerated. Product and Finance estimate $25,000 of incremental weekly contribution margin during the earlier availability window.

That creates an estimated acceleration benefit of $100,000 before incremental AI-program costs. It becomes a realized benefit as the financial outcome is observed and validated.

Do not claim the feature’s entire future revenue as AI’s return. The relevant benefit is the incremental value of the acceleration attributable to the intervention. Some launches only shift bookings between periods; others improve retention or enable sales that would otherwise be lost. The financial model must reflect which occurred.

Use the metrics to improve returns

Once the outcome is clear, investigate its drivers.

Agent Effectiveness combines steering, prompt quality, and task outcomes with measures such as outcome success and cost per outcome. Agent Readiness assesses repository conditions that affect agents’ ability to work. Developer Sentiment captures the friction and confidence that operational telemetry cannot fully explain. [10] [11] [12]

These metrics help choose an intervention. They do not constitute the return themselves.

For a team with expensive failed sessions, inspect the failed work. The cause may be missing environment instructions, weak tests, overly broad tasks, or a poor model choice. Change one bounded part of the workflow and check whether success, delivery cost, and quality improve.

A higher readiness score is evidence that repository conditions changed. Successful delivery at better economics is evidence that the investment paid off.

Use this approach for coaching and system improvement. Team comparisons need context; incident load, ownership, work assignment, and repository conditions can all affect results.

Larridin starting points

Use the metrics to choose an intervention, then check delivery, cost, and quality to see whether it helped.

Put a small scorecard in front of the board

Show the evidence behind the return and the investment decision it supports. Keep estimated capacity separate from verified financial benefits.

Investment

Reconciled AI and implementation costs

Delivery

Output trend with staffing and scope

Priorities

Milestones, allocation, and release evidence

Quality

Mature rework and reliability outcomes

AI contribution

Comparative evidence and limitations

Captured benefit

Verified financial benefit; capacity shown separately

Next decision

What to expand, improve, or stop

Larridin starting points

Pair these measures with release milestones, comparative evidence, and Finance-validated benefits.

Start with a focused 90 day program

01—30

Establish the baseline

Align on business objectives and Engineering’s contribution. Then validate the baseline, spending, data coverage, and quality requirements.

31—60

Evaluate a workflow

Compare like-for-like work. Inspect review and production effects. Record other changes.

61—90

Make the decision

Review with Engineering, Product, and Finance. Separate observations, estimates, and realized benefits.

Start each phase with a small set of metrics.

Keep your chosen measures throughout the program. These are starting points; the wizard tailors the set to your objective.

Days 1–30

Baseline accepted work, AI spending, and involvement. Record staffing, scope, quality thresholds, and data coverage.

Days 31–60

Locate delivery constraints, inspect merge-time risk, and check where additional capacity went. Keep tracking baseline measures.

Days 61–90

Assess economics and quality alongside initiative milestones. Keep immature rework windows provisional; Finance validates captured benefits.

Ninety days is an initial decision window. Continue observing longer-lived work and update provisional results.

Answers worth keeping.

Does higher engineering output prove AI ROI?

No. Higher comparable output can demonstrate additional capacity. Financial ROI requires evidence of incremental business benefit, the investment required, and AI’s contribution to the change.

Can AI Output Share measure productivity improvement?

No. AI Output Share measures involvement under an attribution rule. Estimating incremental impact requires a credible comparison and checks for other explanations.

Are estimated hours saved a financial return?

Estimated capacity value is different from realized financial benefit. Validate the effort estimate, avoid overlapping benefits, and show what the business captured before reporting financial ROI.

How long should an AI ROI evaluation take?

Use 90 days as an initial decision window. Rework, reliability, release outcomes, and business benefits may need longer observation. Mark immature results provisional.

Sources and metric definitions Original research + 12 product references ↗

Product definitions reviewed October 7, 2026. The worked scenarios explain how to apply the framework.

Original research: Larridin Data, August 2026. The benchmark methodology and measurement period are linked alongside the findings.

  1. Engineering Output
  2. Quality and Innovation Rate
  3. Director of Engineering Playbook
  4. Velocity and CI/CD
  5. WorkGraph
  6. Reliability
  7. AI Contribution Metrics
  8. CTO and VP Engineering Playbook
  9. AI Capacity Added
  10. Agent Effectiveness and AI Spend
  11. Agent Readiness
  12. Developer Sentiment
PUT THE FRAMEWORK TO WORK

Build your evidence.
Bring a decision to the board.

Choose your objective, define success and quality guardrails, then turn your results into a board slide and a practical evidence plan.