Skip to main content

Productivity measurement gets harder once humans and AI agents share the same workflows. Login counts and license reports can show that AI is being used, but they don’t show whether a person, an agent, or both produced the work—or whether faster output created more review, rework, or downstream problems.

That distinction matters because perceived productivity and measured productivity don’t always match. In a 2025 randomized controlled trial, METR found that experienced developers took 19% longer to complete real tasks with AI, even though they believed afterward that AI had sped them up by 20%.By May 2026, METR reported that its latest RCT found small productivity benefits from newer AI agents, although the magnitude remained highly uncertain. The takeaway isn’t that AI makes people slower. It’s that organizations need measurement that can distinguish perceived gains from what actually happens in the work.

Larridin CEO Russ Fradin has described the underlying shift as an “atomization of work” between humans and AI agents. Traditional IT monitoring was built to watch people and systems, not the flow of work between people and AI. That visibility blind spot matters: a login report can show that someone used an AI assistant, but it can’t show whether the person did the work, delegated it to an agent, corrected the agent’s output, or worked alongside it throughout the task.

Here are five products worth evaluating for different versions of that problem.

1. Larridin

Larridin is built around the enterprise-wide human-and-agent view. It measures AI adoption and fluency, maps how AI appears inside workflows, connects usage to business outcomes, and brings human and agent spend into the same view. Its browser, desktop, and enterprise data coverage also extends beyond engineering into other departments and workflows.

Best for: CIOs, CFOs, CHROs, and transformation leaders who need to measure AI-powered work across the human and agent workforce, not just inside one technical team.

Limitation: If your main need is code-level attribution between humans and AI inside individual pull requests, an engineering-specific platform such as Weave is more specialized for that job.

2. Weave

Weave (formerly WorkWeave) focuses on engineering intelligence for human and AI work. Its current platform analyzes activity from prompt to production and attributes contributions across pull requests, reviews, and deployments to humans or AI. It also combines AI-specific metrics with DORA, SPACE, surveys, cost, efficiency, and quality signals.

Best for: Engineering leaders who want direct, code-level visibility into how much work humans and AI are contributing and whether AI is improving output, quality, and cost efficiency.

Limitation: Weave is focused on software engineering, so it doesn’t answer the same human-plus-agent productivity question across sales, finance, HR, operations, or other enterprise functions.

3. Faros AI

Faros AI takes a systems-level approach to AI-assisted engineering productivity. Its 2026 Acceleration Whiplash research draws on two years of telemetry from 22,000 developers across more than 4,000 teams, tracking changes in throughput, review work, quality, and production outcomes as AI adoption rises across the software development lifecycle.

Best for: Engineering organizations that want broad telemetry and longitudinal evidence showing how AI changes team and system performance, including downstream effects that raw output metrics can miss.

Limitation: Faros AI is centered on software engineering intelligence rather than productivity across the broader human and agent workforce.

4. GetDX

GetDX combines its developer productivity frameworks with an AI measurement framework focused on utilization, impact, and cost. It uses system metrics alongside developer-reported signals, and its current guidance explicitly recommends measuring autonomous agents as extensions of the developers and teams that oversee them rather than treating agents as isolated contributors.

Best for: Engineering organizations that want human experience, system metrics, AI usage, agent activity, and productivity measurement in one research-backed framework.

Limitation: DX is designed around software engineering and developer experience, so it doesn’t provide the same cross-department view of human and agent work as an enterprise-wide measurement platform.

5. Jellyfish

Jellyfish connects engineering delivery and AI impact to R&D investment and business priorities. Its AI impact capabilities help engineering leaders see who is using which AI tools, how AI affects productivity and code quality, and how engineering time is allocated to AI initiatives.

Best for: CTOs and engineering executives who need AI productivity translated into delivery, capacity, investment, and business-alignment decisions.

Limitation: Jellyfish is oriented toward engineering organization and investment outcomes rather than direct human-versus-agent attribution across the enterprise.

The Real Measurement Problem

The important distinction is what the platform treats as the unit of productivity.

Some tools are strongest at separating human and AI contribution inside engineering work. Others measure how agents affect team delivery, quality, cost, or developer experience. Larridin takes the broader enterprise view, where human and agent work needs to be connected across tools, departments, workflows, spend, and business outcomes.

That makes a few questions especially useful during evaluation: Does the platform measure humans and agents inside the same workflow? Can it separate their contribution when that distinction matters? Does it connect activity to cost and durable outcomes? Does measurement extend beyond engineering? And does the reporting help leaders improve the system rather than simply score individual workers?

The right answer depends on the decision you’re trying to make. If the question is narrowly about AI contribution to code, an engineering-specific platform may be the better fit. If leadership needs to understand how human and agent work is changing productivity across the enterprise, the measurement layer has to be broad enough to follow both.

Talk to an Expert