Skip to main content

OpenAI CFO Sarah Friar says cost per token isn’t enough to evaluate AI investment. She proposed a broader question: how much useful intelligence does each dollar produce? Answering that question requires measuring both the cost of successful outcomes and the value they create.

Key Takeaways

  • Traditional software metrics such as seats, active users, and renewals don’t show whether AI is producing useful work. Friar’s scorecard shifts the focus toward successful outcomes, their full cost, dependability, and whether value improves as usage scales.
  • “Useful Intelligence per Dollar” isn’t a number finance can pull from a provider invoice. Applying it requires connecting AI spend to successful work, quality, retries, human review, and business outcomes.
  • Workflow design is part of the measurement problem. AI can be widely adopted without producing more value if the underlying workflow never changes.

The 4 Questions and What They Actually Require

In a July 2026 OpenAI blog post, CFO Sarah Friar described Useful Intelligence per Dollar as a potential scorecard for the AI era that’s built around four key questions.

1. Is AI completing work that matters?

Friar’s first question moves the discussion from activity to useful output. A login, prompt, or token count can show that AI was used. It doesn’t show whether the work produced something the business needed.

The useful outcome depends on the function. For engineering, it might be accepted code or a merged PR. In customer service, it could be an issue resolved without escalation.

This is where Larridin’s Value and AI Impact measurement matter. The goal is to connect AI activity to delivery, time savings, workflow change, quality, or other outcomes rather than treating usage itself as proof of value.

2. What does each successful task actually cost?

Friar’s second question expands the cost boundary. Token price is only one input. The real cost of a successful task can also include failed attempts, retries, time spent waiting, human review, correction, and rework.

That distinction matters because the cheapest model per token isn’t necessarily the cheapest way to produce an acceptable result. A lower-priced model that needs several attempts and more human correction can cost more per successful outcome than a higher-priced model that gets there with less rework.

Larridin’s Token Spend & Insights provides the AI-spend and attribution side of that calculation across tools, models, agents, teams, and use cases. In engineering, Token Cost Effectiveness can connect token usage, session cost, retries, and accepted outcomes. Organizations still need to account for human review and correction time where those costs materially affect the result.

3. Can people depend on the result?

Dependability is a quality and reliability question. Is the output accurate enough to use? Is it consistent? Does it require correction? Does the workflow know when a human needs to step in?

AI proficiency measures how effectively people use AI, which can influence the reliability of the results. But dependable output also depends on the model, workflow, and task design.

Larridin’s AI Fluency capability can surface proficiency patterns, while AI Impact and outcome data show whether the work is producing reliable, usable results.

4. Does each dollar produce more value as usage grows?

The fourth question is about returns at scale. As AI usage expands, are successful outcomes growing faster than total cost while quality holds or improves? Or is spend simply rising with usage?

Answering that requires tracking performance over time. Leaders need to compare cost, successful outcomes, and quality rather than looking at one month of spend or one adoption snapshot.

Proficiency trends can help explain why the ratio improves: people may learn to choose better tasks, write better prompts, use advanced features, or redesign workflows around AI. But proficiency improvement alone doesn’t prove better economics. The evidence is whether the organization is producing more useful, dependable work for each dollar invested.

How Larridin Helps Operationalize the Scorecard

Larridin’s Utilization × Proficiency × Value framework takes a similar approach, moving AI measurement beyond usage and cost alone to include how effectively people use AI and what value that activity produces.

  • Utilization shows who is using AI, how often, and where it appears in workflows.
  • Proficiency shows how effectively people are using AI and where skills or workflow habits are improving.
  • Token Spend & Insights shows what AI usage costs and attributes that spend across tools, models, agents, teams, and use cases.
  • Value and AI Impact connect AI activity to outcomes such as time saved, delivery changes, quality, and workflow improvement.

Together, those layers let leaders move from “How much AI did we use?” to “What useful work did it produce, what did that work cost, and is the return improving?”

That’s also why workflow redesign matters. McKinsey’s July 2026 research found leaders who redesigned workflows around AI were 5.3x more likely to report enterprise value capture than leaders who left workflows unchanged. Adoption alone doesn’t create useful intelligence. The work has to change in a way the business can measure.

Frequently Asked Questions

Is “Useful Intelligence per Dollar” a measurable metric or a philosophy?

OpenAI presents it as a scorecard rather than a single standardized formula. The four questions provide a way to evaluate whether AI is producing useful work efficiently and reliably as usage grows.

Organizations can turn that scorecard into measurable operating metrics by defining successful outcomes for each workflow, capturing the full cost of producing them, measuring quality and reliability, and tracking the ratio over time.

How does cost per successful task differ from cost per token?

Cost per token is a unit price for model inference. Cost per successful task measures what the organization spent to produce an acceptable outcome.

That can include token consumption, retries, failed attempts, human review, correction, and other work required before the result is usable. It’s possible for a model with cheaper tokens to have a higher cost per successful task.

Where should we start if we can’t answer all four questions today?

Start with one important workflow and define what a successful outcome looks like. Then connect AI usage and spend to that outcome and measure whether quality holds.

That creates the measurement chain needed for the other questions. Once the organization can see successful outcomes and their cost for one workflow, it can expand the same approach to additional teams and use cases.

What does “value at scale” look like in an engineering organization?

It means producing more accepted, durable engineering outcomes per dollar of AI spend while quality and reliability hold or improve. That might show up as more shipped work, shorter cycle times, lower rework, or better cost per accepted outcome.

The important part is the combination. More AI usage by itself isn’t value at scale, and lower cost by itself isn’t either.

Build the Measurement Infrastructure Behind the Scorecard

Useful Intelligence per Dollar changes the AI budget conversation from how much the organization used to what that usage produced.

Larridin connects utilization, proficiency, spend, and impact so leaders can measure useful work, understand what successful outcomes cost, and see whether AI value is improving as adoption scales.

Book a discovery call to see how Larridin connects AI spend to successful outcomes and business value.