AI Compute Costs Fall 60% a Year, but Enterprise Bills Keep Rising Under Jevons Paradox

AI Compute Costs Fall 60% a Year, but Enterprise Bills Keep Rising Under Jevons Paradox

N
News Editor 01
2026-07-24 05:15:17
Anthropic CEO Dario Amodei says AI inference costs are falling by about 60% annually, yet Red Hat's Brian Gracely argues enterprise spending keeps climbing as cheaper compute drives more usage.

AI inference costs are falling by about 60% a year, but enterprise bills are still moving in the opposite direction. Speaking at a VentureBeat AI event, Red Hat portfolio strategy director Brian Gracely said companies have moved past the experiment phase and are now facing a harsher question: where is the return on investment, and how do they stop AI spending from running ahead of budgets?

Gracely described this shift as an AI “Day 2 moment.” Getting a model to run in a sandbox is one thing. Running it in production with stability, governance, and cost discipline is much harder. In his view, cost control, governance, and long-term sustainability are proving tougher than the initial buildout itself.

Cheaper inference can still lead to bigger total bills

He tied the pattern to Jevons paradox, the economic idea that efficiency gains do not always reduce consumption. In some cases, they do the opposite. Applied to AI, lower per-inference costs can trigger a much larger jump in usage, leaving total spending higher even as unit economics improve.

Gracely said some customers have bought 50,000 Copilot licenses without a clear view of what employees are actually getting from them. What they do know, he said, is that they are paying for expensive GPU-backed compute. The problem is no longer whether AI can be deployed. It is whether the spending can be measured and controlled.

From token consumption to workload control

His proposed response is for enterprises to stop thinking only as passive token consumers and start acting as active token managers, or even token producers in some cases. That means asking which workloads truly require the newest top-tier models and which can be handled by open-source or smaller models.

Internal knowledge search, document summaries, and common customer service queries may not need the most expensive model available. Keeping premium models for tasks that actually demand deeper reasoning can improve the cost structure. Gracely said that could mean operating GPUs directly or renting them, but the central issue is gaining more control over important workloads.

Open-source models are changing the buying equation

The market is also giving enterprises more room to choose. Two years ago, options were concentrated among a small group of cloud providers and packaged services such as Copilot. Now, with the rise of open-source models including DeepSeek, companies have more technical alternatives and more leverage in pricing discussions.

Gracely ended with a reminder that AI has really only been around for three years in its current wave and is still in an early stage. For enterprises, the immediate task is not chasing scale for its own sake. It is identifying which AI licenses and compute expenses are producing measurable value and which ones are simply adding cost.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
1300

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.