Skip to content
Agent Effectiveness · part of Developer Intelligence

See how your engineers work with agents. Make the best way the default.

Agent Effectiveness analyzes coding-agent sessions to help teams understand which practices produce useful, verifiable work. Compare session patterns, review coaching suggestions, and share practices across the team.

Claude CodeCodexCursorOpenCodeGitHub Copilot

Spotlights

Example sessions · Last 30 days

74 / 100

team effectiveness, example top quartile 81

12 of 64

engineers already scoring 80 or above

Plan before you build

19 engineers open feature work with a plan request. 34% fewer iterations, 6 example prompts.

+4 Feature development$1,900 / mo

Adoption 19 of 64 · discuss as a team practice

Share

Name the files before the agent edits

Diffs 2.2× smaller when the prompt names the edit surface. 23% of feature sessions do.

+3 Feature development

Adoption 14 of 64

Share

Run the suite before reporting done

56 of 97 bug-fixing sessions already do. The other 41 needed 2.9× more follow-up fixes.

Suggested CLAUDE.md practice+5 Bug fixing

Adoption 56 of 97 sessions · review the guide change

Review

Illustrative practice suggestions. Confirm access and attribution settings with your organization's administrator.

Trusted by AI-forward enterprises

  • Vertiv
  • OK! magazine
  • Gainsight
  • The Signatry
  • Klaviyo
  • SurveyMonkey
  • Globality
  • TigerConnect
  • ConnectPay
  • The Joint Chiropractic
  • EcoVadis
  • Rev.io
  • Belcorp
  • Sundt
  • Polk County, WI
  • Source Advisors
  • Andelyn Biosciences
  • University of Hertfordshire

Agent mix

See the work behind your agent usage and cost.

Larridin brings coding-agent sessions, work patterns, token usage, and cost into one view. Session cost estimates use model pricing and usage telemetry; provider billing views show billed spend where a billing connection is available.

Work patterns · compare feature development, bug fixing, research, testing, and other work in the illustrated view
Session context · inspect agent configuration and linked pull requests alongside the session
Prompt cache accounted · cache creation and cache reads have different prices; inspect their contribution to estimated cost
Coverage matters · interpret comparisons alongside the sessions and coding tools available in the selected period

Agent mix

Example · Last 30 days · Engineering · 64 engineers

$21,480

estimated cost of analyzed sessions

612

analyzed sessions of 1,529 captured

$35

average per session, median $4

28%

of estimated cost in six sessions

Daily estimated cost by type of work

weekends lighter

$1,800
$1,200
$600
Aug 10
Aug 17
Aug 24
Aug 31
Feature developmentCodebase researchBug fixing & debuggingPrototypingTesting & qualityOther (ops, data, docs)
Type of workShareSessionsScoreSpend
Feature development11876$10,310
Bug fixing & debugging9766$3,437
Codebase research9679$2,148
Prototyping1971$1,933
Testing & quality16281$1,504
Ops & maintenance5870$859
Docs & knowledge3580$645
Data & analytics2773$644
Six sessions account for $6,014. The other 606 average $26. Review the six first.

Prompt cache

input tokens, all agents

91%

of input tokens read from cache in this example. Cache share varies across sessions.

$9,240

estimated difference from uncached input pricing in this example.

$412

estimated cache-creation cost across 62 example sessions.

Illustrative cost estimates use model list rates, including separate cache-creation and cache-read prices. Estimates can differ from invoices. Scores summarize the analyzed sessions in each group.

Sessions

Open any session. See what the agent did, what it cost, and how the pair worked.

Review available session traces, agent turns, token usage, estimated cost, and linked PRs. The score combines six dimensions: prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes. Available dimensions carry equal weight.

Prompts, duration, cost and tokens · review time and usage signals alongside the session; the example shows elapsed and engaged time separately
A score with citations · inspect the practices and evidence behind a session score; the example below shows selected dimensions
Connected agents in one list · compare Claude Code, Codex, Cursor, and other connected agents with their linked PRs
How the session score works

Agent Traces

App: All
All · 1,529 Merged PR · 80 Needs attention · 166
ScoreSessionAppPRsEnded ↓
53Document router-prompt scoringClaude Code1Aug 31
77Larridin brand voice guidelinesClaude Code0Aug 30
57Optimize dbt CI pipelineClaude Code0Aug 29
81Git worktree setup automationClaude Code0Aug 29
67HubSpot lead generation syncClaude Code0Aug 29
70AI Impact dashboard filtersClaude Code0Aug 29
60GCP dbt job log analysisClaude Code0Aug 29
63Plugin and authentication flowClaude Code0Aug 28
57Local persistence for draftsClaude Code0Aug 28
67Spend Mapping UI migrationCodex0Aug 28
37Performance impact of the routerClaude Code0Aug 28
70AI Spend Intelligence exportCursor0Aug 27

Prompts

32

30h 23m open

Engaged

2h 40m

4 re-warms

Cost estimate

$37.54

Example model mix

Tokens

76.5M

99.5% cached

EvaluationRecommendationsDetails

Overall score

81 / 100 ▲ 9 vs team's August average of 72

A strong session. The engineer opened with an explicit goal and asked for a plan, held scope with five targeted corrections, and the agent verified its work before reporting completion.

Show full reasoning →

16 citations · 9 prompts, 7 agent turns

Selected dimensions · 3 of 6

team average

Prompt clarity77

Asks are scoped with an explicit session goal, constraints and an upfront plan request.

Prompt quality82

Each prompt combines the goal, a performance target and constraints for the CI speed-ups.

Session steering91

Catches drift quickly with targeted corrections and clear scope boundaries.

Scoring

Find useful practices and give your team guidance they can act on.

Session scores bring six equally weighted dimensions into a team view of agent use. Compare practices, inspect the sessions behind a pattern, and turn useful findings into guidance your team can review and adopt.

Adoption, outcomes and cost per practice · compare how a practice relates to outcomes in the sessions available to your team
Compare practices across your own sessions · look for patterns within comparable work and consistent reporting periods
Reviewable guidance · use a suggested practice to draft an agent guide change, review it with your team, and track later sessions

Effectiveness

Example · Last 30 days · Engineering

74 / 100

▲ 2 vs August average of 72

12 of 64

engineers scoring 80 or above

6

dimensions in the effectiveness score

Example median · 68 Example top quartile · 81 your team · 74

Working well

Plan before you build

19 engineers · feature sessions finish in 34% fewer iterations

82

Name the files before the agent edits

23% of feature sessions · diffs 2.2× smaller

79

Targeted corrections instead of restarts

41 engineers · steering at top-quartile level

79

Tests run before reporting done

56 of 97 bug-fixing sessions · the fewest follow-up fixes in the org

83

Needs attention

Closing without a test run

41 of 97 bug-fixing sessions · 2.9× more follow-up fixes

58

Opening with "fix the bug"

33 sessions · 12 turns before a plan appears

61

Restarting instead of steering

27 sessions · a fresh session after each miss, context rebuilt every time

64

Idle gaps that rebuild the cache

62 sessions · $412 in cache rebuilds, no score effect

$412

The score averages available ratings for prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes, scaled to 100. Comparisons here are illustrative.

Example practice guidance

review with your team

1

Run the test suite before reporting done

41 sessions closed without one, then needed 2.9× more follow-up fixes.

CLAUDE.md · apps/api+5 Bug fixing$640 / mo

Review
2

Ask for a plan before feature work starts

19 of 64 engineers already do. Share their pattern with the other 45.

Spotlight+4 Feature development$1,900 / mo

Discuss
3

Compare models for codebase research

96 research sessions. A lower-cost model scored 78 against 79 at 42% less in this example.

Model comparison$900 / mo

Explore

Example guide change · apps/api/CLAUDE.md

## Before you report done+ Run `pnpm test --filter api`, paste the summary line.+ If a test fails, fix it or name it and say why.+ Never report done on a red suite.
Review suggested text Discuss with team

For you

example coaching

Six of your nine bug-fixing sessions this month closed without a test run. Teammates who run the suite first finish with 2.9× fewer follow-ups. Try /verify on the next one.

Example coaching based on session patterns. Access depends on your organization's roles and configuration.

Privacy

Make access clear from the start.

Agent sessions can contain sensitive context. Larridin scopes analytics by organization and role. Work with your administrator to confirm which session details are collected, who can access them, and the retention settings for your deployment.

Confirm session access · agree who needs session detail and who should use team summaries
Review collection settings · check collector and redaction settings before enabling transcript collection
Use team summaries · review team scores, work mix, and spend; confirm detailed access with your administrator
Trust Center · SOC 2 Type II · GDPR

Access planning example

Recommended configuration

EngineerManagerAdmin
Transcript and promptsOwn sessionsPer policyPer policy
Session scoresOwn sessionsTeam rollupOrg rollup
Work mix and spendOwn sessionsTeam, by engineerOrg
Coaching suggestionsPersonalTeam guidanceOrg guidance

This table illustrates an access policy to review with your administrator. It does not describe default permissions; available controls depend on your deployment.

A closer look

How the Agent Effectiveness Score works

Each scored session is evaluated across six skills. The composite is the mean of the available 1–5 dimension scores, converted to a 0–100 scale. Unobserved dimensions are excluded rather than treated as zero.

Prompt clarity and quality

Assess how clearly the engineer states the goal and supplies useful context and constraints.

Steering and sophistication

Assess how the engineer guides the session and uses the agent’s capabilities to work through the task.

Verification and outcomes

Assess whether the work is checked and what the session accomplished. Follow the session evidence behind the score.

Coding agents represented in the product’s tool catalog
Coding agentsWhat to check before comparing
Claude Code, Codex, Cursor, OpenCodeConfirm session capture and attribution for the versions your organization uses.
GitHub Copilot, Devin, Gemini CLI, AntigravityA tool appearing in the catalog does not imply identical session coverage. Review captured and unscored sessions.

Questions about Agent Effectiveness

What does the Agent Effectiveness Score measure?

It measures observed skills in a coding-agent session: prompt clarity, session steering, sophistication, prompt quality, verification discipline, and task outcomes. It is not an overall rating of an engineer’s productivity.

How should we compare scores?

Compare sessions with similar work types and coverage using the same rubric. Read the evidence and number of scored sessions alongside any average. Differences can reflect task difficulty, capture coverage, or work mix.

Who can see transcripts and session data?

Session data is organization-scoped and access depends on role, enabled tools, and sharing configuration. Review access controls with your administrator before capturing sensitive sessions. Model-provider retention and Larridin session storage are separate controls.

How should I interpret the product examples?

The product views on this page use illustrative data to explain the metrics and workflows. For metric definitions, sample calculations, and assumptions, see our measurement methodology.

Start this week.

Connect your coding agents and review the sessions available to your team. Try it on your own repos, or get pricing for your org.