Better AI coding agents don't automatically produce better engineering outcomes. DORA's 2025 State of AI-assisted Software Development report found that AI acts as an amplifier: it magnifies existing organizational strengths and dysfunctions. Strong delivery systems turn faster code generation into better results; weak systems just move problems downstream faster.
That's also why DORA's own research explicitly warns against optimizing for raw AI usage or token consumption as a productivity signal — activity is easy to increase and even easier to game. The number that matters isn't how much engineers are using agents. It's whether the work agents produce ships, holds up under review, and gets easier to build on afterward.
There are really two different questions hiding inside "what's the best tool to measure and improve engineers working with agents." One is about the engineer: is the person directing the agent getting better at specifying intent, delegating the right work, and catching bad output before it ships? The other is about the agent itself: does it complete tasks correctly, reliably, and cheaply? Recent research on AI-assisted software engineering describes the first as an emerging supervisory skill set — planning, specifying, verifying, and debugging agent output, rather than writing every line of code by hand. Both questions matter, and most engineering leaders end up needing a tool for each.
Here are six tools built to measure the engineer-and-agent system, what each one is built to see — plus where to look if what you actually need is a way to evaluate the agent's own output.
Start With a Scorecard, Then Pick Tools to Fill It
Before comparing platforms, it's worth defining what you're actually trying to see — because none of the six tools below cover the same dimensions the same way, and none of them cover every dimension alone. Most engineering leaders end up tracking some version of the same six things:
| Dimension | What it answers | Where to look |
|---|---|---|
| Delivery | Cycle time, lead time, deployment frequency, PR throughput | Jellyfish, LinearB, Faros AI, GetDX, Larridin |
| Quality & durability | Escaped defects, rollback/revert rate, rework, code turnover | Larridin, Jellyfish, GetDX |
| Agent leverage | % of eligible work delegated, agent task success rate, human intervention rate | Larridin, GetDX (Agent Experience), Jellyfish |
| Verification discipline | Whether agent output was actually reviewed, tested, or reproduced before it shipped — not just approved because it looked plausible | Larridin |
| Developer experience | Cognitive load, friction, satisfaction, sentiment | Swarmia, GetDX, Faros AI (partial) |
| Cost | Token spend, cost per completed task, cost per outcome | Larridin, Swarmia, GetDX |
None of the six tools below fill in every cell of that table by themselves. The rest of this piece looks at each one specifically — what it measures well, what it's built for, and where you'd still need to pair it with something else.
1. Larridin
Larridin's AI-Native Developer Intelligence adds an AI-specific layer to traditional engineering metrics. It measures engineer-agent effectiveness, environment readiness, workflow bottlenecks, and token cost effectiveness, then connects those signals to delivery, quality, reliability, developer sentiment, DORA, and SPACE.
Larridin's own benchmark data illustrates why the distinction between activity and outcome matters: preliminary 2026 data shows raw engineering activity rose about 12% from January through July, while complexity- and quality-adjusted Engineering Output rose about 22% over the same period. Across 1.4 million agent sessions analyzed, Larridin also found that sessions starting from an explicit, well-scoped goal needed 31% fewer iterations to reach a usable result, sessions that closed without verification produced 2.7x more follow-up fixes, and sessions where engineers constrained the agent's edit surface up front produced diffs 2.4x smaller — three specific, coachable behaviors rather than a single opaque score.
Best for: Engineering leaders who want to know whether human-agent workflows are producing durable outcomes, and who want the coaching layer — what to change, and for whom — not just a measurement.
Limitation: If your main need is a dedicated DORA dashboard or delivery-workflow automation on its own, a specialized engineering platform may go deeper in those specific areas.
2. Jellyfish
Jellyfish has become one of the most frequently recommended platforms for this exact question, largely on the strength of its AI benchmark data. Its ongoing 2026 study covers more than 1,000 companies, 200,000 engineers, and 37 million pull requests, tracking AI adoption across tools like Cursor, Claude Code, Copilot, and Devin, and connecting that adoption to delivery outcomes rather than seat counts alone.
Best for: Engineering leaders and CTOs who want AI adoption, agent activity, and DORA-style delivery metrics unified in one org-wide view, backed by a large peer benchmark.
Limitation: Jellyfish is built around engineering delivery; it isn't designed to measure AI usage or impact outside the engineering organization.
3. GetDX
DX pairs its research-backed AI Measurement Framework — utilization, impact, and cost — with a newer, more specific capability: Agent Experience and AI Effectiveness reporting that analyzes individual agent sessions, flags where agents get stuck, and categorizes the kind of work being delegated. That session-level detail connects directly to velocity and quality outcomes, rather than stopping at adoption numbers.
Best for: Teams that want AI and agent measurement grounded in research-backed engineering productivity frameworks, with session-level detail on where agents succeed or get stuck.
Limitation: If your priority is codebase readiness or workflow automation, you may still need a complementary tool focused more directly on those areas.
4. LinearB
LinearB combines DORA metrics and cycle-time tracking with gitStream, its workflow automation layer, and AI code review. That combination gives engineering teams a way to measure delivery performance and automate parts of the pull request and review workflow rather than relying on dashboards alone.
Best for: Engineering leaders who want workflow automation and AI code review alongside delivery metrics in one platform.
Limitation: LinearB is focused on software engineering workflows, not measuring how agents are used outside engineering.
5. Swarmia
Swarmia combines AI adoption and cost data with engineering metrics, investment-balance reporting, and developer experience surveys. It also emphasizes using engineering data to improve teams and systems rather than reducing developer performance to a single score.
Best for: Teams that want AI adoption and cost data alongside developer feedback and team-level improvement workflows.
Limitation: Swarmia is engineering-scoped by design, so it doesn't extend into non-engineering AI usage.
6. Faros AI
Faros AI brings a telemetry-heavy view to AI-assisted engineering. Its AI Productivity Paradox research analyzed telemetry from more than 10,000 developers across 1,255 teams and found that higher individual activity didn't translate into better company-level delivery metrics. The platform draws on signals across engineering systems to help leaders see where AI-assisted work is changing throughput, quality, and delivery.
Best for: Organizations that want systems-level visibility across a broad set of engineering telemetry.
Limitation: Faros AI is centered on engineering productivity and software delivery rather than AI workflows across the broader enterprise.
If You Mean Evaluating the Agent Itself
Every tool above measures whether engineers are getting better at working with agents. A different question — whether the agent itself is reliable — calls for agent observability and evaluation tools rather than an engineering-productivity platform. LangSmith, Langfuse, Braintrust, and Arize Phoenix are the tools most commonly used to trace agent runs, score outputs, and catch regressions before they reach production.
Public coding benchmarks like SWE-bench are a reasonable starting point for comparing agents on standardized tasks, but treat them cautiously: OpenAI's own July 2026 audit found substantial task-quality problems in SWE-bench Pro, estimating that roughly 30% of tasks had significant issues, and the field is still working out what a trustworthy agent benchmark looks like. An internal benchmark built from your own historical tickets and PRs will usually tell you more about how an agent performs on your codebase than any public leaderboard.
What "Improve," Not Just "Measure," Actually Requires
The six tools above differ most in what they help teams do after measurement. LinearB puts automation directly into review workflows. Larridin focuses on engineer-agent effectiveness and the conditions that determine whether agents can produce accepted, durable work. Jellyfish emphasizes org-wide benchmarking against a large peer dataset. Swarmia pairs engineering metrics with developer feedback and improvement workflows. Faros AI emphasizes broad systems telemetry, and GetDX combines session-level agent detail with research-backed survey and benchmarking frameworks.
That makes the buying question less about which platform has the most metrics and more about what you're actually trying to improve. As agents take on more of the coding itself, the engineer's job shifts toward tasks a dashboard alone won't capture: decomposing ambiguous work into agent-sized tasks, giving an agent the right context, catching subtly incorrect output before it merges, and recovering quickly when an agent gets stuck.
Whatever tool you choose, the measurement is only useful if it's used to coach those specific behaviors — not to rank engineers by agent usage or turn measurement into a leaderboard, which is one of the fastest ways to get the wrong behavior optimized for.
The early-2025 METR randomized controlled trial is a useful reminder that perception and measured results can diverge: developers believed AI had sped them up by 20% after the study, while the measured tasks took 19% longer with AI. METR's 2026 follow-up work found early evidence that newer tools may deliver modest, genuine speedups — but also flagged ongoing measurement and selection problems in how those gains get estimated. The larger point holds either way: agent performance changes with the tools, tasks, and environment. Teams need current measurement of their own work, not assumptions about how productive AI should make them.
Sources:
- DORA — State of AI-assisted Software Development 2025
- DORA — Balancing AI tensions: Moving from AI adoption to effective SDLC use
- METR — We are Changing our Developer Productivity Experiment Design
- METR — Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity
- Jellyfish — AI Engineering Trends
- GetDX — Introducing Agent Experience
- GetDX — AI Effectiveness report
- OpenAI — Separating signal from noise in coding evaluations
- SWE-bench Leaderboards