Most engineering teams treat human code as the quality baseline and AI code as the unknown. One Larridin customer’s 30- and 90-day revert data challenges that assumption and shows why quality needs to be measured by code source.
The difference in a Larridin customer’s revert rate data is significant, so it needs to be framed carefully. Over 30 days, AI-assisted code had a 0.2% revert rate, while human-only code had a 15.16% rate. Human-only code reverted roughly 76 times as often.
The pattern held over 90 days, although the size of the gap changed. AI-assisted code reverted at 0.94%, compared with 27.8% for human-only code—a roughly 30x difference.
This doesn’t prove AI writes better code than humans in every setting. It means that in this environment, across these workflows and measurement windows, AI-assisted commits were substantially less likely to be reverted. The data surfaces the difference without explaining it. Better code quality, narrower task selection, different review practices, or some combination could all contribute.
The broader code turnover picture helps leaders test whether low revert rates reflect durable work or leave other forms of rework uncounted.
METR’s 2025 study has a useful warning about relying on self-assessment. In a randomized trial of 16 experienced open-source developers working in mature repositories, participants took 19% longer when AI tools were allowed. Afterward, they estimated that AI made them 20% faster.
METR now says those results are out of date and no longer reflect the current impact of AI models on open-source developer productivity. That limits what the study says about today’s tools. But it doesn’t erase the measurement lesson: perceived performance and observed performance can differ sharply in a specific environment.
The same discipline applies to AI code quality. “AI code is sloppy” and “AI code is better” are both opinions until repository data shows whose work is reverted, rewritten, or retained.
Revert rate measures how often committed code is undone within a defined window. A low rate means the code was less likely to require a direct reversal during that period.
That makes revert rate a useful indicator of near-term code durability, especially in environments where reverts are the standard response to a problematic change.
A low revert rate doesn’t prove the code is high quality. Code can stay in the repository while accumulating technical debt, introducing subtle regressions, or creating an architectural problem that takes longer to surface.
The 30- and 90-day windows measure near-term durability. A more complete quality view also needs code turnover, rework, incidents, test behavior, and longer-term maintainability.
A single aggregate revert rate can’t show whether durability problems are mainly in AI-assisted or human-only work. The measurement system needs to distinguish code source at the commit or pull request level, then apply the same time windows to each group.
Larridin’s AI Dev Productivity platform puts AI code share and code durability in the same view, giving engineering leaders a comparison aggregate metrics can’t provide.
Revert rate and code turnover give engineering leaders a concrete starting point. Ask four questions:
Those questions produce more useful evidence than a debate based on individual anecdotes. Larridin’s 5-dimension engineering measurement framework places code durability alongside throughput, workflow performance, quality, and cost.
It depends on the environment, workflow, and measurement window. The Larridin customer data shows a large revert-rate advantage for AI-assisted code in one instance. It’s a real result from that environment, not a universal rule. The useful question is what the same comparison shows inside your organization.
The best starting benchmark is internal. Compare AI-assisted and human-only code in the same repositories over the same period using the same definition. The customer data showed 0.2% versus 15.16% at 30 days, but the important signal is how your own rates compare and change.
It means the code was durable within the measurement window. That matters, but it isn’t enough on its own. Low-revert code can still create technical debt, subtle defects, or long-term maintenance problems. Pair revert rate with turnover, rework, incidents, and other quality signals.
Connect version-control data with AI tool telemetry or other reliable source attribution, then classify commits or pull requests as AI-assisted or human-only. Apply the same 30- and 90-day windows to both groups so the comparison stays consistent.
Larridin tracks revert rates, code turnover, and 30- and 90-day durability by code source. This is the data leaders need to replace gut instinct with an actual quality comparison.
Book a discovery call to see your code durability picture.