AI spend is getting harder to forecast with a seat-count spreadsheet.
The FinOps Foundation says predictability is generally lower for AI costs, especially in earlier-stage deployments, where pricing and consumption can be inconsistent and forecasts may need to combine cloud and non-cloud cost components.
EY points to another source of uncertainty. Agentic AI shifts enterprise spending toward variable compute consumption, and EY notes that current pricing may understate the long-term economics if upstream providers are absorbing or subsidizing some of the underlying compute cost.
Traditional software forecasts often start with seats, renewal prices, and headcount. AI adds more variable cost drivers.
Some costs are still fixed subscription fees. Others change with token use, API calls, model selection, agent activity, credits, and usage-based charges. As AI systems become more agentic, that variable portion can become more important.
There’s another problem: billed spend doesn’t always tell you how much AI was actually consumed. Provider credits, discounts, bundled allowances, and subsidies can temporarily reduce an invoice even when usage is high. Our AI token spend attribution guidance recommends looking at billed spend and observed usage together when building forecasts for different pricing, credit, and usage scenarios.
A single organization-wide average can hide another source of forecast error. In one Larridin customer environment, one engineer generated 65% of the team’s AI spend in one week, while several teammates spent between $0 and $300. Forecasting everyone from the same per-person average would flatten the pattern that actually drove the bill.
The budget is what the organization plans to spend. The forecast is an estimate of where spending is headed.
That difference matters because AI usage is changing quickly. A quarterly or annual budget can stay fixed while the forecast moves because a team starts using agents more heavily, changes models, adds a new workflow, or consumes credits faster than expected.
A useful forecast should be updated when the underlying behavior changes. Its job is to give finance and engineering enough warning to decide whether the new trajectory is expected, valuable, and affordable before the budget period closes.
Don’t build the forecast from seat price alone.
Track what providers bill and, where available, the underlying token, model, request, or agent usage that created the charge. The two numbers answer different questions: billed spend tells finance what the organization owes now, while observed usage helps show how costs could move when credits, discounts, pricing, or workload mix changes.
Our Token Spend & Insights consolidates billed spend and usage across providers and attributes it to teams, agents, projects, workflows, models, and use cases.
Don’t assume every user or workload behaves like the average.
Separate predictable subscription costs from variable consumption, then segment variable spend where the data supports it: by team, tool, model, agent, project, or use case.
Spend concentration deserves special attention. If a small number of users or workflows account for a large share of consumption, model those drivers separately instead of spreading their behavior across the whole licensed population.
This also makes changes in the forecast easier to explain. A rise in total spend can be traced to a specific team, model, agent, or workload instead of being described as generic AI growth.
A single point estimate can look more certain than the data deserves.
Build a base case from current usage and known commitments, then show what could move the number higher or lower. Useful assumptions might include adoption growth, heavier agent usage, a shift in model mix, different provider credits, or a new high-volume workflow.
Pricing itself should also be treated as a variable. EY notes that current agentic AI pricing may not fully reflect the underlying compute economics and that providers are already moving from subscription pricing toward consumption-based models.
The important part is to make the assumptions visible. A CFO can evaluate “our forecast rises if agent usage doubles” much more easily than a precise annual number with no explanation of what could change it.
A forecast shouldn’t stay fixed until the next planning cycle.
Update it as actual spend and usage arrive. Compare forecast to actuals, investigate meaningful variance, and revise the assumptions when the pattern changes.
Our Token Spend & Insights projects spend by team and flags budgets at risk before the quarter closes. That turns forecasting into an early-warning system instead of an explanation for why the invoice missed the plan.
Longer-term forecasts have another source of uncertainty: no one cost trajectory is guaranteed.
A 2026 RIKEN analysis examines several possible AI market scenarios through 2030 as memory costs, inference efficiency, open models, infrastructure economics, and competition change. Some point to persistent premium pricing, while others suggest much stronger commoditization.
The important forecasting lesson is that today’s pricing shouldn’t be treated as a fixed assumption several years into the future.
For multi-year planning, create more than one pricing scenario and identify which parts of the budget are most exposed to a change. A workload that depends heavily on frontier models, for example, has a different risk profile from one that can shift among lower-cost models as economics change.
That approach also keeps the forecast useful when the market moves in a direction nobody predicted exactly.
Use enough recent data to represent normal activity for the teams and workloads you’re forecasting. There’s no universal 60- or 90-day rule.
For a new deployment, start with the data you have and use a rolling baseline that becomes more reliable as usage develops. For an established deployment, make sure the baseline includes enough history to capture normal variation rather than one unusually light or heavy week.
Keep the underlying cost logic by provider or tool rather than forcing everything into one token rate.
Track billed spending alongside observed usage where available, and account for subscriptions, usage charges, credits, discounts, and other pricing terms separately. That lets the forecast change when the pricing structure changes without rebuilding the entire model.
It needs input from both, with a clear owner for the forecast.
Finance or FinOps understands budgets, commitments, invoices, and forecast methodology. Engineering understands why usage is changing: which teams are adopting new workflows, where agents are being introduced, and which model or tool changes are planned.
The forecast is stronger when both are working from the same attribution data.
Present a range and show the assumptions behind it.
Explain the base case, the variables most likely to push spending higher or lower, and which signals you’re watching. That’s more useful than hiding uncertainty behind a precise number that may be obsolete as soon as usage changes.
Model more than one plausible pricing path.
The RIKEN analysis illustrates why: changes in hardware economics, model efficiency, open-model competition, and infrastructure markets can produce very different outcomes. Rather than betting the forecast on one prediction, show what happens under higher-cost, lower-cost, and relatively stable-cost scenarios.
AI spend forecasting gets more credible when finance can see both the bill and the behavior behind it.
Our Token Spend & Insights connects billed spend and observed usage to the teams, agents, projects, workflows, models, and use cases driving cost, then projects spend and flags budgets at risk before the quarter closes.
Book a discovery call to build your AI spend forecast from real consumption data.