Why Your AI Budget Numbers Are Probably Wrong
🏢 University of Oxford
📅 August 2026
Here is the most useful number I have seen in months: 87% of the real decline in AI prices is invisible to the measurement methods most companies and statistical agencies still use. That means the tools you are comparing suppliers with are not showing you how much cheaper the capability has actually gotten.
The paper behind it, “The Price of Intelligence,” joins 21,024 posted prices across 3,208 models and 86 providers to 4,605 benchmark scores. Louis Yiven Zhu at Oxford built what economists call a quality-adjusted price index, the same logic statistical agencies apply when they strip out improvement to track how much a television or a laptop really costs. He did it for AI, and the result should change how every finance and procurement team budgets.
![]()
Why This Paper Matters
Every executive has been told AI gets cheaper each year. It does. But the official-style numbers badly understate how much. When Silicon Valley Certification Hub talks to finance leaders, the same mistake comes up again and again: people budget AI off sticker token prices and assume the market is barely moving. The real question is what a capability costs, once you account for the massive jump in quality.
That distinction is not academic. It changes competition and concentration. If the capability is getting seven times cheaper than reported, a market that looks settled and expensive is actually dynamic and cheap. Decisions about which vendors to go with, whether to build or buy, and how fast to move all shift when you see the true curve.
Methodology, Explained Simply
Think of how you judge whether your phone got cheaper. A new model costs the same as last year’s, but it takes better photos and runs faster, so in quality-adjusted terms it actually got cheaper. Economists call this a hedonic comparison. The trick with AI is that “quality” has to come from somewhere concrete.
Zhu built his quality ladder from benchmark scores, 4,605 of them, and estimated which models sit where on the ladder from how they respond on those tests. Then he did the comparison twice. One method, matched-model, is what agencies use for software: only compare a model to itself over time, ignore new models entirely. The other, quality-adjusted, lets new and better models enter and tracks the price per unit of real capability.
Here is where the counterintuitive kicker lives. When you count cost per completed task rather than per token, the buyer’s price stopped falling. Reasoners now chew through far more tokens to finish a job than the drop in token price can offset. The seller’s price and the buyer’s price diverge, and the gap is a trap waiting for anyone who budgets per token.
![]()
Results and Practical Insights
Kick with the 87%. Under valid, pre-registered measurement, AI capability is roughly seven times cheaper than the official numbers suggest. The result survived a robustness audit: even when the six most disputed benchmarks were dropped, the steep decline persisted and widened. So this is not a fragile result built on one test.
At the same time, per completed task, the price to you flattened. Cheap tokens hid a rising real cost as models reasoned longer. For a Chief AI Officer running an AI Assessment for companies, the lesson is blunt: do not buy intelligence by the token, and do not report savings by the token either. Measure the cost of a completed outcome, and re-measure often, because the curve moves fast.
Are you budgeting your AI on sticker token prices or on the real cost per completed outcome?
At Silicon Valley Certification Hub, we help finance, procurement, and operations leaders evaluate and deploy AI that fits their actual business processes, with economics measured on the results that matter.
What This Means for Your Chief AI Officer
Silicon Valley Certification Hub keeps telling executives that governance is only half the job. The other half is measuring whether the investment is working, and this paper shows most teams are measuring it wrong. Treat the official price indices and your own sticker invoices as the floor, not the truth.
Your real lever is an AI Assessment for companies that ties spend to completed business outcomes. If a model now finishes a task your team needed split across three prompts a year ago, your per-outcome cost fell even when the invoice looks flat. And if a reasoning model quietly triples its tokens to get there, your real cost can climb while the headline price falls. Both things can be true at once, which is why per-token budgeting fails.
Key Takeaways for Finance and Operations Leaders


Thanks to All Authors
Want to know how this applies to your company?
At Silicon Valley Certification Hub, we help you align AI + Strategy. Our team works directly with your directors and teams to assess AI readiness, identify gaps, and build a clear path forward — tailored to your business context.
Book a time with our CEO, Alejandro Cuauhtemoc-Mejia:
https://calendar.app.google/2ihQf2JH3D9uJBe68
Silicon Valley Certification Hub
3000 El Camino Real, Building 4, Palo Alto, CA
0 Comments