AI Is Getting Cheaper to Build. It's Getting More Expensive to Run.

The AI cost curve everyone's watching is going down. The one that actually hits the P&L is going up.

Quick answer

  1. The cost per token collapsed a 280-fold drop in about eighteen months yet global AI spending is on pace to hit $2.59 trillion in 2026. Price was never the cost driver; usage is.
  2. AI has no free marginal user: every call costs real money, which is why AI-native businesses run thinner margins (62% vs. the 80% SaaS norm) and agents make it worse, at 5–30x the tokens of a chatbot exchange.
  3. Fixing it is an operating discipline across four teams Finance, Operations, Engineering, and Sales each owning its piece at the same time.

Key Takeaways to remember:

  1. Price fell, bills rose: cheaper tokens didn't lower anyone's bill usage is growing faster than price is falling.
  2. AI isn't SaaS economics: there's no near-free hundredth user, so margins run lower than the software businesses AI is replacing.
  3. Agents multiply the cost: 5–30x the tokens of a single chatbot call, and adoption is climbing fast the bill only goes up.
  4. The real risk is success: Uber burned its entire annual AI tooling budget in four months, then capped usage at $1,500/month per employee, per tool.
  5. The fix needs four owners at once: Finance, Operations, Engineering, and Sales none of them works in isolation.

Building an AI feature has never been easier. A prototype that used to take a quarter now takes a weekend, sometimes an afternoon. Model access is a credit card, not a research budget.

Running that same AI at scale is a different story, and it's the one almost nobody has budgeted for.

Here's the number every CFO should already know: the cost of running a GPT-3.5-level model fell from $20 per million tokens in late 2022 to $0.07 by late 2024, according to Stanford's AI Index. That's a 280-fold drop in about eighteen months. Now put that next to this: global AI spending is on pace to hit $2.59 trillion in 2026, up nearly 50% from last year, per Gartner.

So the price of AI collapsed, and companies are spending more on it than ever. That's not a contradiction. It means the price was never the real cost driver. Usage is. And usage is growing faster than price is falling.

The price was never the real cost driver. Usage is. And usage is growing faster than price is falling.

The Economics Nobody Re-Modeled

Here's why that catches most companies off guard. Traditional software makes money because the hundredth user costs almost nothing extra to serve. Once you've built the product, adding a customer is nearly free. That's why the median software subscription business still runs an 80% gross margin today, per 2025 benchmark data from Aleph and Benchmarkit.

AI doesn't work that way. Every time someone uses it, it costs you real money, whether that's the first user or the millionth. There's no free marginal user. The same benchmark shows usage-based pricing running at 62% gross margin, eighteen points lower than the software norm, and that gap is a big part of why AI-native companies already run thinner margins than the SaaS businesses they're replacing.

And spending hasn't slowed down to compensate. Enterprise GenAI spend tripled in a single year, from $11.5 billion in 2024 to $37 billion in 2025, per Menlo Ventures. Cheaper tokens didn't lower anyone's bill. They just made it affordable to use AI for more things, more often.

Why Agents Make It Worse

Agents push this further, and the reason is simple. A chatbot answers one question and stops one model call, done. An agent doesn't stop there. It plans the task, calls tools, checks its own work, and retries if something's wrong, which can add up to 10 to 20 model calls to finish one job a person used to just do themselves. Gartner's own analysis puts agentic workloads at 5 to 30 times the tokens of a single chatbot exchange.

Now multiply that by how fast companies are adopting agents. Gartner's 2026 forecast for AI agent software alone runs to nearly $207 billion, up 139% from $86.4 billion the year before. That's the trap: the more useful agents get, the more they cost to run, and there's no version of this where adoption goes up and the bill goes down.

The Real Risk Is Success

Here's what that looks like inside a real company. Uber rolled out agentic coding tools to its engineers and put usage on an internal leaderboard to encourage adoption. It worked, maybe too well. By April, four months into 2026, Uber's CTO confirmed the company had already burned through its entire annual AI tooling budget. The fix was a hard cap: $1,500 a month, per employee, per tool.

Nothing about that story is a failure. The tools worked exactly as intended, and people used them exactly as encouraged. That's what makes it the real risk. Nobody had built a budget for a pilot that worked this well, this fast, because most AI budgets are still sized for adoption that trickles in, not adoption that takes off.

What Actually Has to Change

Fixing this isn't one team's job. It splits cleanly across four, and it only works if each one actually owns its piece.

1. Finance: Rebuild How the Budget Is Structured

Separate "cost to build" from "cost to run" as two different lines, not one. Building has a start and an end. Running scales with adoption, indefinitely, and needs to be forecast the way you'd forecast revenue rather than a license renewal. Only 26% of enterprises say they have real-time visibility into what their AI actually costs to run, per KPMG's Q2 2026 AI Pulse survey. That's a finance problem before it's anything else, and it's the one most companies haven't touched yet.

2. Operations: Govern the Rollout, Not Just the Invoice

Set a consumption ceiling per team, per tool, before a pilot scales company-wide not after it's already spent, the way Uber found out the hard way. And name one owner for the number. Accountability for AI spend currently splits almost evenly between finance and technology, 53% and 55% respectively, according to DoiT's own research. Split accountability works like no accountability. Someone needs the authority to hold up a rollout that doesn't have a cost model behind it.

3. Engineering: Architect for Cost the Way You Architect for Capability

Route simple requests to cheap models. Cache what repeats instead of recomputing it every time. Save the expensive, frontier-level model for the handful of calls that genuinely need that much reasoning. Most enterprise AI spend today runs the expensive model on every request by default, simply because nobody built the routing layer that would make the cheap model the default instead. This is the one lever that cuts the bill without touching a single budget line or contract.

4. Sales & Commercial: Price for What It Costs to Serve

A flat seat price sitting on top of a variable-cost product isn't a pricing strategy. It's a margin problem with a delay timer attached. If the cost of running your product scales with usage, the price to the customer should scale with it too even if that means a harder conversation with sales about how deals get structured.

None of these four work in isolation. Finance can build the cleanest forecast in the world, and it won't matter if engineering hasn't built the routing layer to act on it. The point isn't picking one. It's making sure all four have an owner, at the same time.

The Stance

Put all of this together and the competitive question changes. It used to be who ships the best AI feature first. It's becoming who can run AI profitably at the scale customers actually want to use it.

That's not a model decision anymore. It's an operating discipline, and it belongs on the same agenda as the product roadmap, not buried in a line item under cloud. The companies that build that discipline now will be the ones still standing when everyone else is explaining their invoice.

FAQ: The Cost of Running AI at Scale

If AI got 280x cheaper, why are bills going up?

>

Because price was never the real cost driver usage is. Cheaper tokens didn't lower anyone's bill; they made it affordable to use AI for more things, more often. Usage is growing faster than price is falling, which is why global AI spend is still on pace for $2.59 trillion in 2026.

Why don't AI products have SaaS-like margins?

>

Traditional software has a near-free hundredth user, which is how the median subscription business runs an 80% gross margin. AI has no free marginal user every call costs real money so usage-based pricing runs closer to 62% gross margin, about eighteen points lower.

Why do agents cost more to run than chatbots?

>

A chatbot is one model call and done. An agent plans, calls tools, checks its own work, and retries 10 to 20 calls to finish one job. Gartner puts agentic workloads at 5 to 30 times the tokens of a single chatbot exchange, so the more useful agents get, the more they cost.

Haresh KumbhaniCTO

Haresh Kumbhani leads Zymr’s solution architecture and technology strategy. A hands-on technical leader and serial entrepreneur, Haresh brings decades of complex product development and deployment experience.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Ready to get started?

Book a free strategy call or explore more guides