Engineering Measurement
Published engineering metrics so far: agent effectiveness and team performance.
Agent Effectiveness
- Agent Effectiveness Score How do we measure whether our engineers are using AI coding agents effectively, and where do we need to invest in training rather than better tools?
- Outcome Success What percentage of AI coding agent sessions in our engineering organization actually result in code that gets merged, and what's driving the sessions that don't?
- Cost per Outcome How do we calculate the true cost per unit of engineering output from AI coding agents, and how does that cost vary across the different tools our team uses?
- Session Depth How deeply are our engineers engaging with AI coding agents in each session, and what does session depth tell us about proficiency versus unproductive looping?
- Spend How does the amount we're billed for AI coding agents compare to what our telemetry shows we actually consumed, and what does the gap tell us about our agent cost governance?
- App Adoption How do we compare the cost-efficiency of every AI coding agent we're running, across adoption, sessions, spend, and outcomes, in a single view?
- Model Spend Breakdown Which foundation models are our AI coding agents calling, how much is each one costing, and are we routing to expensive frontier models for tasks that cheaper models handle equally well?
- AI Fluency Score Summary What specific engineering behaviors and infrastructure elements make up our AI Fluency Score, and which ones should we invest in developing to improve agent effectiveness?
- Team Breakdown How do we see each engineer's AI coding agent proficiency and cost efficiency side by side, so we can identify who to learn from and who needs development support?
Team Performance
- Output per Engineer How do we measure whether AI coding tools are actually increasing the engineering output our team produces, rather than just increasing the number of commits or PRs?
- Output by Category What proportion of our engineering team's AI-assisted output is going toward new features versus maintenance work, and is AI investment changing that mix over time?
- AI Output Share / AI PR Share Is AI coding producing a proportionate share of our engineering team's most complex and valuable work, or is it handling the simpler tasks while humans carry the heavier PRs?
- Prompt Quality How good are our engineers at prompting AI coding tools, and how much productivity are we leaving on the table because of gaps in prompting skill rather than gaps in tool capability?
- Cost per Output Is our AI coding tool investment making our engineering team more cost-efficient over time, or are costs rising as fast as, or faster than, the productivity gains?