The moment AI usage becomes a number next to an engineer's name in a review, that number stops measuring what you think it measures.
Key Takeaways
- AI usage and output metrics can help diagnose engineering workflows, but they’re poor measures of individual performance.
- Individual comparisons miss context such as task type, codebase complexity, and review practices, so the same AI-usage number can mean very different things.
- Use team, repo, and work-type data to find workflow problems. Individual data is better suited to coaching, enablement, and knowledge sharing than ranking.
Rankings Change the Behavior You're Measuring
AI leverage measures the gap between AI capacity added and AI capacity actually used. It shouldn't be used to rank individual engineers. Tying AI metrics to performance reviews can encourage engineers to increase visible activity, such as prompts or generated code, whether or not that activity improves engineering outcomes .
The Same AI-Usage Number Can Mean Different Things
Consider two teams with almost identical AI usage. Both have 20 engineers, about 90% active AI users, roughly 400 agent sessions a month, and similar token spend.
The outcomes are very different. One team gives agents narrow, well-scoped tasks and verifies the output before review; about 260 sessions become merged, durable work. The other gives agents broader tasks and accepts larger generated diffs with less scrutiny; only about 90 sessions become merged work, and more of that work is later reverted or rewritten.
The same problem applies to individual comparisons. AI usage can vary because of ticket complexity, codebase conditions, task mix, or review discipline. A developer with high AI usage may have found a better workflow. Or they may simply be working on more automatable tasks.
Engineering output makes the same point: an accepted AI-assisted pull request can reflect good task scoping, strong tests, and an experienced engineer as much as the model itself. A per-person ranking can't separate those factors.
What to Measure Instead
The goal is to measure AI-assisted engineering at a level where the context is useful. Track:
- Adoption: Where AI is being used across teams, repos, and work types.
- Usage: How much AI activity is occurring, without treating activity alone as productivity.
- AI leverage: How much of the AI capacity you've added is actually being used.
- Engineering outcomes: Whether AI-assisted work is accepted, durable, or later reverted or rewritten.
Larridin’s AI-native developer intelligence focuses on diagnosing workflows rather than ranking engineers. The useful question is what engineering outcome a given amount of AI usage or spend produced, and where the delivery system is helping or getting in the way.
Use Individual Data for Coaching, Not Ranking
If your organization already has individual-level AI data, it can still be useful. It may help identify engineers who have found effective AI workflows, teams that need more enablement, or people who could benefit from help with a specific tool or task.
The difference is what happens next. Using the data to start a coaching or knowledge-sharing conversation preserves the context around the number. Putting the same number into a leaderboard and tying it to performance reviews turns the number into the conclusion.
Frequently Asked Questions
Should AI usage be part of an individual performance review?
Not as a standalone productivity score or ranking. AI usage doesn't account for task complexity, codebase conditions, review practices, or the quality and durability of the resulting work.
Can managers use individual AI data at all?
Yes. Individual data can help with coaching, training, and identifying useful practices that can be shared across a team. The risk comes from treating a usage or output metric as a direct measure of individual performance.
Measure AI Engineering Performance Without Individual Leaderboards
Larridin helps engineering leaders connect AI adoption, usage, spend, and engineering outcomes at the team and workflow level without turning AI metrics into individual leaderboards.