Skip to main content

EY illustrates how a customer-service interaction can rise from $0.04 to $1.20 when a linear AI workflow becomes an orchestrated agentic system. The lesson is that the architecture behind one interaction can radically change its economics.

Key Takeaways

  • Agentic workflows can cost far more per interaction than simpler AI systems because each interaction can involve multiple model calls, tools, reasoning steps, and other compute-intensive work.
  • Agentic workloads can vary dramatically even when performing the same task, and higher token consumption doesn’t necessarily produce better results.
  • Total token spend isn’t enough to judge agent economics. Track what a successful workflow costs, how often it succeeds, how much human intervention it requires, and what value the result produces.

Why the 30x Number Matters

In EY’s example, the higher cost comes from the additional work happening behind a single interaction. The agentic workflow adds tools, reasoning, planning, subagents, and iterative loops to what was previously a straightforward input-retrieval-response workflow. Each additional step can add model and infrastructure usage, increasing the total cost of the interaction.

That changes what finance needs to measure. Total monthly spend shows what the organization paid, but not whether the agent delivered enough value for that cost.

For customer service, a more useful measure might be the cost of a successfully resolved interaction. For engineering, it might be the cost of a completed change that meets the team’s quality requirements. The right unit depends on the outcome the agent is supposed to produce.

Why Agentic Costs Can Vary So Much

Agentic work isn’t just more expensive than a single model interaction. It can also be much less predictable from one run to another.

A 2026 study of eight frontier models performing agentic coding tasks found that those tasks consumed about 1,000x more tokens than code reasoning and code chat. Input tokens drove most of the overall consumption.

The study also found that separate runs on the same task could differ by as much as 30x in total token use. Higher consumption didn’t reliably improve accuracy. In many cases, performance peaked at an intermediate level of token use and then stopped improving as consumption increased.

That creates a different cost problem from a fixed software license. Two attempts at the same type of work can produce very different bills even when the intended outcome hasn’t changed.

Context accumulation is one reason agentic workflows can consume so many input tokens. Our guide to context accumulation and agent costs looks more closely at that mechanism.

The Token Bill Is Only Part of Agent Cost

Token and API charges are the most visible expenses, but they aren’t the whole operating picture.

EY separates agent cost into seven categories. Six cover current operating costs: tokens and API calls, subscriptions and licenses, platform infrastructure, governance, organizational change, and expected failure and recovery. A seventh category covers potential future AI-specific taxes and is explicitly identified as speculative.

Those costs can land in different budgets. Model usage appears on one invoice. Infrastructure may appear on the cloud bill. Human review can show up as engineering or operations time. Failure recovery may not appear as a recognizable AI cost until something goes wrong.

That makes the vendor invoice an incomplete basis for comparing agents with the work they replace.

For a useful comparison, include the costs required to run the workflow and produce an acceptable result.

4 Metrics for Evaluating Agent Cost and Value

1. Cost per Successful Workflow

Start with what the agent is supposed to complete.

Define a successful outcome for the use case, then calculate what it costs to produce one. A completed run that fails the quality threshold shouldn’t count the same as one that delivers the intended result.

This gives finance a more useful number than cost per model call or aggregate monthly tokens because it connects spending to completed work.

2. Success and Human-Intervention Rates

Measure how often the agent completes the work without failing, restarting, or requiring someone to step in.

Human oversight is sometimes part of the intended workflow. The important thing is to account for it rather than treating the agent’s token cost as the full cost of completion.

If one agent looks inexpensive but repeatedly sends work to a human for correction, the apparent savings can disappear when the complete workflow is measured.

3. Cost Variation Across Runs

An average can hide expensive outliers.

The frontier model study found up to 30x variation in token consumption across runs of the same coding task. That makes the distribution useful alongside the average.

Track whether similar work tends to stay within a predictable cost range or whether a small number of runs account for a disproportionate share of spending. That distinction matters when deciding whether an agent is ready to scale.

4. Value per Successful Outcome

Higher agent cost isn’t automatically bad.

A more expensive workflow can still make sense if it completes valuable work faster, reduces human effort, increases throughput, lowers risk, or produces another measurable benefit worth more than the additional cost.

That’s why EY argues that agent spending should be managed as an investment tied to measurable return rather than evaluated on cost alone.

The useful comparison is agent cost against the value of the completed work.

Separate Agent Spend Before You Evaluate It

That analysis gets much harder when human AI use and autonomous agent activity are mixed together.

A developer using an AI coding assistant has a different consumption pattern from an agent that can keep executing without a person initiating every step. Combining both into one AI spend total hides those differences.

Our Token Spend & Insights separates human-driven and agent-driven spend and attributes costs to teams, agents, workflows, and use cases. That gives finance and technology leaders a way to compare what different agentic workflows cost instead of judging them from one aggregate token bill.

Once agent spend is separated, teams can track cost per workflow alongside completion, quality, and business outcomes and decide which uses earn additional investment.

Frequently Asked Questions

Does EY’s 30x example mean every agentic workflow costs 30x more?

No. EY compares a simple customer-service workflow costing $0.04 per interaction with a more complex orchestrated version costing $1.20. It illustrates how architecture can change unit cost; it isn’t a universal multiplier that applies to every agent or use case.

Organizations need to measure their own workflows to determine the actual difference.

Why does agentic AI generally consume more than a linear AI interaction?

Agents can perform multiple steps before returning a final result. Those steps can include planning, retrieval, tool calls, reasoning, subagent activity, validation, and additional model requests.

Should we measure agent cost per interaction or per task?

Use the unit that best represents the business outcome.

For some customer-service applications, an interaction or resolved case may make sense. For engineering, finance, or operational agents, a completed workflow, issue, transaction, or other successful outcome may be more meaningful.

The important thing is to connect cost to completed value rather than choose a metric only because it’s easy to count.

Can we use average cost per agent run?

Yes, but don’t rely on the average alone. Agent costs can vary substantially between runs. Track the average alongside the range and the expensive outliers so one apparently reasonable number doesn’t hide a small number of unusually costly executions.

How do we know whether an expensive agent is worth keeping?

Compare its full cost with the value of the successful outcomes it produces. Include model usage, relevant infrastructure, human oversight, and failure or recovery costs. Then compare that total with the time saved, work completed, revenue supported, risk reduced, or other business result the agent is expected to deliver.

An expensive agent can still have strong ROI. A cheap agent that consistently fails can still be poor value.

Measure What Agents Cost and What They Produce

Agentic AI changes cost structures by making spend more variable and more dependent on how a workflow executes.

The answer isn’t to assume every agent is too expensive or to judge the deployment from the aggregate token bill. It’s to measure what each workflow consumes, whether it succeeds, and what the successful result is worth.

Our Token Spend & Insights separates agent spend from human AI use and connects consumption to the agents, workflows, and use cases driving it.

Book a discovery call to see what your agents cost and what that spending produces.