A client called me two weeks ago sounding like he’d found cash in an old coat pocket.
He’d been reading about the AI price war. New providers undercutting the incumbents, model costs falling every quarter, the number he was quoted in March suddenly looking like the wrong number. His question was simple: should he push his project to Q4 and let prices drop further?
I told him no. Not because prices will stop falling, they won’t. Because he was staring at the smallest line on the invoice.
The price collapse is real. It just isn’t the story.
The numbers are not in dispute. Per-token pricing for a given level of capability has fallen as much as 280-fold since 2022, according to figures Gartner and the Stanford AI Index both track. July brought another round. Meta opened its first paid model API on July 9 and priced below the flagship offerings from OpenAI and Anthropic. Everybody else answered.
If you buy AI the way you buy electricity, this looks like great news. Same power, lower rate.
Except that’s not how any of this works.
Here’s the number that should actually get your attention: while unit prices were collapsing, enterprise AI budgets were climbing. Not holding steady. Climbing. In April, Uber’s engineering organization burned through its entire annual AI coding budget with two-thirds of the year still ahead of it. Two months later Microsoft pulled back most internal developer access to a coding assistant for the same reason — cost. Neither of those companies is careless with money, and neither of them is short on engineering talent.
They got caught by the same math that’s about to catch a lot of businesses much smaller than they are.
Cheaper tokens, bigger bills
The math is not complicated once you see it.
A chatbot query is one round trip. You ask, it answers, you pay for two small piles of tokens. That’s the workload everyone priced their expectations against in 2024.
An agent is not that. An agent plans, retrieves, calls a tool, checks its own work, hits an error, retries, and calls another tool. Gartner’s data puts agentic workflows at five to thirty times the token consumption of a simple query for the same user-visible action. The FinOps Foundation reported companies running three times over their annual token budgets on agentic workloads specifically.
So the rate dropped by an order of magnitude and the meter started spinning by two. That’s the whole story of 2026 in one sentence.
It’s compounding, not stabilizing. Goldman Sachs projects token consumption climbing toward twenty-four times current levels by 2030. Falling unit prices are not going to outrun that curve.
The costs that never appear on the rate card
Even that understates it, because the rate card only covers inference.
In every engagement I’ve run, the supporting infrastructure has cost more than the model calls. The data pipelines that feed it. The vector store and what it costs to keep it current. The egress charges nobody models in advance. The logging and observability you need to answer “why did it do that.” The evaluation harness that tells you whether last week’s model update quietly broke something.
None of that is on the price sheet you’re comparing across vendors. All of it is on your invoice.
Then there’s the cost that doesn’t show up as a bill at all. Workday’s research found that for every ten hours AI saves a team, employees spend close to four hours fixing its output. That’s not a rounding error. That’s forty percent of your headline productivity gain walking back out the door. It lands in payroll rather than in your cloud spend, which is exactly why most business cases miss it.
What the price war actually proved
Set the invoices aside for a second, because there’s a bigger point buried in all of this.
If model costs fell 280-fold and outcomes didn’t improve, then the model was never what was standing between you and results.
The evidence backs that up hard. Writer’s 2026 enterprise survey found only 29% of organizations reporting significant ROI from generative AI, and just 23% from agents. Those numbers did not move meaningfully while prices cratered. Cheap intelligence, in the hands of a company with tangled data and no integration layer, produces cheap disappointment faster.
I watched the opposite version of this at Kyndryl on one of my accounts. We drove a 38% improvement in incident detection and documented more than $2M in savings. The model was not the expensive part of that work, and it was not the hard part either. The hard part was the eighteen months of unglamorous foundation work underneath it. Getting the telemetry clean, getting the systems talking, getting 200+ engineers to actually change how they worked on a Tuesday morning.
Drop the model cost to zero and none of that gets easier. Which means none of it gets cheaper.
How to budget for AI when the model is nearly free
So what do you actually do with a collapsing price curve? Four things.
Meter the workflow, not the token. Your unit of cost is not a million tokens. It’s one completed business outcome: one invoice reconciled, one ticket resolved, one report produced. Instrument that from day one. If you can’t tell me what a single resolved ticket costs you in AI spend, you cannot manage this line item, and no price cut will save you.
Split infrastructure spend from AI spend in the business case. These are two different investments with two different lifespans. Your data pipelines and integration work outlive whatever model you’re using this quarter, and they hold their value when the model gets swapped. Blending them into one number is how you end up in front of a CFO defending “AI” as a line item when most of the money went to plumbing you needed regardless.
Set your cost-per-outcome ceiling before you write the first line of code. What is a resolved ticket worth to you? Whatever that number is, your ceiling sits underneath it. Decide it while you’re calm and skeptical, not six months in when you’ve got sunk cost and a stakeholder who wants the win.
Build a kill threshold. Uber didn’t have one. That’s the actual lesson, and I wrote about the deeper version of it in what really went wrong with their AI budget. A threshold is a number that, when crossed, forces a review. Not an apology, a review. Without one, agentic spend has no natural stopping point, because every individual call looks reasonable in isolation.
The question the price war is really asking you
“Wait for prices to drop” stopped being a strategy the moment it failed at Uber and Microsoft. Both of them had cheaper tokens than they’d ever had. Both blew the budget anyway.
Which leaves you with a decision, and it’s binary.
You can treat the falling price curve as permission to move faster on the same shaky foundation: more agents, more calls, more spend, against data you already know isn’t ready. Cheap tokens will let you scale that mistake at a speed that would have been financially impossible two years ago.
Or you can treat it as a window. The model is no longer the constraint or the cost center. That frees up budget and attention for the work that actually determines whether any of this returns a dollar: the data, the integration, the security posture, the adoption.
One of those paths ends with a surprise invoice and a stalled project. The other ends with a system that works.
You’re going to pick one this year, whether you decide to or not.
Ready to find out what your AI is actually going to cost?
Start with an AI Infrastructure Assessment. Before you commit budget to a model, get a real picture of what’s underneath it — infrastructure readiness scorecard, data pipeline analysis, integration gaps, and an executive briefing with honest ROI risk analysis. Get in touch.
See how we work. Our services run infrastructure-first, in sequence, with no shortcuts — assessment, strategy, implementation, and the change management that makes adoption stick.
Keep reading. More on AI budgets, vendor selection, and the foundation work nobody puts in a press release, on the Summit AI blog.
Russell Love is the Founder & CEO of Summit AI Business Solutions, based in Browns Summit, NC. With 20+ years of enterprise transformation experience at IBM and Kyndryl, Russell helps businesses build the foundations that make AI actually work.