Companies bringing AI into their operations are facing a growing cost problem. Most AI services are priced by tokens, but the quality and length of generated output remain hard to predict, leaving businesses with limited visibility over operating costs. Pricing structures for large language models also change frequently, squeezing third-party software vendors at the same time that AI agent usage is rising fast.
That combination has turned cost control into a central issue for the technology sector: how to weigh the value of AI tools against a billing model that can shift quickly and scale unexpectedly.
Why token budgets are hard to lock in
Large language models break prompts and responses into tokens, the basic unit used for billing. The challenge is that AI systems are non-deterministic. The same prompt may not produce the same output every time, and even modest changes in instructions or switching to another model can materially change token usage.
For enterprises, that makes it difficult to establish a fixed cost model over a two- to three-year period. The problem gets larger when companies begin linking multiple AI agents together to make collaborative decisions, because both token use and uncertainty can rise sharply in that setup.
Lower unit costs have not solved the spending problem
Token costs have fallen in recent years, but overall usage is moving in the opposite direction. Goldman Sachs forecast that as enterprises shift heavily toward AI agents, global token consumption will increase 24-fold between 2026 and 2030, reaching 120 quadrillion tokens per month.
Many companies do not discover how far spending has drifted until the monthly invoice arrives. ABMedia cited Uber as an example, saying the company used up its annual token budget in just a few months.
Deploying an AI agent inside a product can be extremely easy. Management may only need to click a button, a much faster process than changing headcount plans or hiring staff. But that convenience carries hidden costs. Security testing and building safeguards also consume tokens, and those expenses can rise quickly without drawing much attention at the start.
Frequent upstream price changes complicate product pricing
AI music company Sumo said it is struggling to decide how to charge customers. The company has considered broad price increases, outcome-based charging, and per-item pricing. But when upstream LLM providers adjust prices every few months, the underlying pricing structure becomes unstable.
Sumo also cannot easily pass those costs on to users, according to the report, because customers do not like variable pricing.
Multiple AI agents can drive costs higher
In modern AI systems, a single model often cannot meet more complex needs on its own. Companies are relying more heavily on several AI agents working together to execute tasks. In that arrangement, chained calls between agents and more frequent interactions can push token consumption even higher.
How companies are trying to control token burn
Businesses are using several practical methods to build cost guardrails around AI deployments.
Restricting expensive tools
Some companies directly limit the use of specific external AI tools by employees or engineers. Microsoft, for example, has restricted the use of certain third-party AI coding tools among its engineers to reduce unnecessary spending.
Demanding more precise prompts
Companies are also training staff to give AI systems more exact and detailed instructions. The report compared this to sending a family member shopping with a clear grocery list. Better prompts can reduce irrelevant or excessive responses and cut wasted token usage.
Choosing models more carefully
Rather than chasing the newest or most complex model, companies are being urged to evaluate which system best fits a specific task and offers a more reasonable cost-performance balance.
Budgeting for testing and safeguards in advance
Many product managers count development costs when rolling out AI, but security protections and system testing also require substantial token use. Companies designing those controls need to account for that spending upfront.
Using personal accounts as a temporary workaround
Some smaller businesses are using personal accounts to avoid more expensive enterprise traffic pricing. But experts cited in the report warned that, as major AI companies face growing pressure from shareholders to generate profits, that gray area is likely to face broader cleanup. Companies should not treat it as a long-term answer to cost control.

