Skip to content

Enterprise AI Measurement Guide

Agent Effectiveness

Model Spend Breakdown

Which foundation models are our AI coding agents calling, how much is each one costing, and are we routing to expensive frontier models for tasks that cheaper models handle equally well?

What it shows

The Model Spend Breakdown shows how agent spend is distributed across the underlying AI models being consumed, not just the coding apps, but the foundation models (Claude Sonnet, Claude Opus, GPT-4o, Codex, and others) that those apps call. This view reveals whether spend is concentrating in frontier models or distributing across tiers, and which models are driving the majority of agent cost in the engineering org.

Why it matters

Model selection is the most impactful cost lever in AI coding, and most engineering organizations do not have a coherent model routing strategy. If Opus is handling tasks that Sonnet would handle equally well, the organization is paying 3-4x the per-prompt cost without a corresponding quality gain. The Model Spend Breakdown is the starting point for model routing decisions: it shows where frontier model spend is concentrated and where cheaper tiers are already handling comparable workloads.

The Larridin angle

Model Spend Breakdown scoped to agent activity is particularly important because agentic workflows consume tokens at 10-50x the rate of chat interactions, due to context accumulation across turns. A model that looks cheap at the chat layer becomes the dominant cost driver at the agent layer, a distinction only visible when spend is tracked at the model level within agent sessions.

Related Agent Effectiveness Metrics

See how your organization measures up

Larridin turns every metric in this guide into a live, benchmarked dashboard for your org. No spreadsheets, no manual surveys.