OpenAI CFO Sarah Friar says cost per token isn’t enough to evaluate AI investment. She proposed a broader question: how much useful intelligence does each dollar produce? Answering that question requires measuring both the cost of successful outcomes and the value they create.
In a July 2026 OpenAI blog post, CFO Sarah Friar described Useful Intelligence per Dollar as a potential scorecard for the AI era that’s built around four key questions.
Friar’s first question moves the discussion from activity to useful output. A login, prompt, or token count can show that AI was used. It doesn’t show whether the work produced something the business needed.
The useful outcome depends on the function. For engineering, it might be accepted code or a merged PR. In customer service, it could be an issue resolved without escalation.
This is where Larridin’s Value and AI Impact measurement matter. The goal is to connect AI activity to delivery, time savings, workflow change, quality, or other outcomes rather than treating usage itself as proof of value.
Friar’s second question expands the cost boundary. Token price is only one input. The real cost of a successful task can also include failed attempts, retries, time spent waiting, human review, correction, and rework.
That distinction matters because the cheapest model per token isn’t necessarily the cheapest way to produce an acceptable result. A lower-priced model that needs several attempts and more human correction can cost more per successful outcome than a higher-priced model that gets there with less rework.
Larridin’s Token Spend & Insights provides the AI-spend and attribution side of that calculation across tools, models, agents, teams, and use cases. In engineering, Token Cost Effectiveness can connect token usage, session cost, retries, and accepted outcomes. Organizations still need to account for human review and correction time where those costs materially affect the result.
Dependability is a quality and reliability question. Is the output accurate enough to use? Is it consistent? Does it require correction? Does the workflow know when a human needs to step in?
AI proficiency measures how effectively people use AI, which can influence the reliability of the results. But dependable output also depends on the model, workflow, and task design.
Larridin’s AI Fluency capability can surface proficiency patterns, while AI Impact and outcome data show whether the work is producing reliable, usable results.
The fourth question is about returns at scale. As AI usage expands, are successful outcomes growing faster than total cost while quality holds or improves? Or is spend simply rising with usage?
Answering that requires tracking performance over time. Leaders need to compare cost, successful outcomes, and quality rather than looking at one month of spend or one adoption snapshot.
Proficiency trends can help explain why the ratio improves: people may learn to choose better tasks, write better prompts, use advanced features, or redesign workflows around AI. But proficiency improvement alone doesn’t prove better economics. The evidence is whether the organization is producing more useful, dependable work for each dollar invested.
Larridin’s Utilization × Proficiency × Value framework takes a similar approach, moving AI measurement beyond usage and cost alone to include how effectively people use AI and what value that activity produces.
Together, those layers let leaders move from “How much AI did we use?” to “What useful work did it produce, what did that work cost, and is the return improving?”
That’s also why workflow redesign matters. McKinsey’s July 2026 research found leaders who redesigned workflows around AI were 5.3x more likely to report enterprise value capture than leaders who left workflows unchanged. Adoption alone doesn’t create useful intelligence. The work has to change in a way the business can measure.
OpenAI presents it as a scorecard rather than a single standardized formula. The four questions provide a way to evaluate whether AI is producing useful work efficiently and reliably as usage grows.
Organizations can turn that scorecard into measurable operating metrics by defining successful outcomes for each workflow, capturing the full cost of producing them, measuring quality and reliability, and tracking the ratio over time.
Cost per token is a unit price for model inference. Cost per successful task measures what the organization spent to produce an acceptable outcome.
That can include token consumption, retries, failed attempts, human review, correction, and other work required before the result is usable. It’s possible for a model with cheaper tokens to have a higher cost per successful task.
Start with one important workflow and define what a successful outcome looks like. Then connect AI usage and spend to that outcome and measure whether quality holds.
That creates the measurement chain needed for the other questions. Once the organization can see successful outcomes and their cost for one workflow, it can expand the same approach to additional teams and use cases.
It means producing more accepted, durable engineering outcomes per dollar of AI spend while quality and reliability hold or improve. That might show up as more shipped work, shorter cycle times, lower rework, or better cost per accepted outcome.
The important part is the combination. More AI usage by itself isn’t value at scale, and lower cost by itself isn’t either.
Useful Intelligence per Dollar changes the AI budget conversation from how much the organization used to what that usage produced.
Larridin connects utilization, proficiency, spend, and impact so leaders can measure useful work, understand what successful outcomes cost, and see whether AI value is improving as adoption scales.
Book a discovery call to see how Larridin connects AI spend to successful outcomes and business value.