The CFO sees the AI bill. The CTO sees adoption. Neither number answers the question they both need answered: What does each AI tool cost to produce a unit of shipped engineering work? That’s what cost per shipped outcome measures.
The value of cost per shipped outcome becomes clear when two tools have similar adoption but very different costs per accepted outcome. In one Larridin customer environment, Codex and Claude Code both had 91% weekly active adoption across 10 of 11 engineers. But the observed AI spend per merged PR was very different:
That’s a 2.7x gap despite identical adoption. It tells finance and engineering where to investigate. Are the tools being used for comparable work? Are some tasks more complex? Does one require more retries or abandoned sessions? Does the higher-cost tool produce better-quality or more durable work?
The metric identifies the gap. The surrounding engineering data helps explain whether it matters.
The basic calculation is straightforward:
AI tool spend for the measurement period ÷ shipped outcomes attributed to that tool = cost per shipped outcome
For an engineering team, a shipped outcome might be a merged PR, completed task, or another unit of accepted work that can be connected to the AI activity that helped produce it.
Use the AI spend associated with the tool during the same period as the outcomes you’re measuring.
That should include usage from sessions that didn’t produce accepted work. Retries, abandoned sessions, and exploratory work still consumed AI resources even when nothing ultimately shipped.
Larridin’s Token Spend & Insights consolidates AI spend across tools and models and attributes it to teams, agents, projects, and use cases.
For coding agents, merged PRs are one useful option because they move the comparison beyond raw AI activity and into work the engineering system accepted.
They aren’t the only option. Larridin’s Engineer-Agent Effectiveness framework also points to accepted diffs, resolved tasks, verified changes, and durable artifacts as possible engineering outcomes.
The important part is to use the same outcome definition when comparing tools.
A lower cost per shipped outcome means less observed AI spend was associated with each accepted outcome during that period. It does not automatically mean the tool delivered more value.
Work complexity, repository conditions, review requirements, output quality, and later rework can all affect the comparison.
That’s why Larridin’s developer measurement approach pairs accepted outcomes with durability and quality signals rather than treating a merged PR as the end of the story. Engineer-Agent Effectiveness specifically distinguishes raw usage from accepted, durable engineering work.
OpenAI CFO Sarah Friar’s Useful Intelligence per Dollar scorecard asks a broader question: What does each successful task actually cost?
For a business, that calculation includes AI usage as well as employee time, human review, retries, and rework. Larridin’s cost per shipped outcome captures the AI-spend portion of that calculation for an accepted engineering result. It does not represent the full cost of producing that result.
Cost per session tells you what each AI session costs on average. Cost per shipped outcome connects total AI spend with accepted engineering work, including spend from sessions that never produced an accepted result.
No, it means the AI spend per accepted outcome was lower. Tool comparisons should also account for work complexity, quality, durability, and rework before making an investment decision.
You need enough comparable outcomes to keep one unusual task or week from dominating the result, and compare tools over the same period and similar types of work. Track the ratio over time rather than treating one snapshot as a verdict.
Internal comparisons are usually more useful because codebases, work mix, engineering practices, and review standards differ. Compare tools within the same environment first, then use external benchmarks as context rather than a pass/fail threshold.
Larridin connects AI usage and spend with accepted engineering outcomes so finance and engineering can see what different tools cost relative to the work they produce.
Book a discovery call to see how your AI coding spend connects to shipped outcomes.