The AI market suffered another violent correction this week, triggered by a revelation from Meta. The company, which previously touted a massive self-built compute infrastructure to fuel superintelligence, now signals it may sell its excess AI compute capacity to third parties. The immediate market reaction assumed the $100 billion+ AI capex boom was about to implode. However, a closer examination reveals a far more nuanced story.

Almost simultaneously, Palantir CEO Alex Karp appeared on CNBC to deliver a blistering critique of the token-based pricing model used by frontier labs like OpenAI and Anthropic. Karp claimed CEOs are privately complaining that they are paying for 'tokens that create zero value' while handing over their proprietary data. He even described the increasingly expensive model bills as a 'wealth tax' on enterprises. This accusation lays bare a growing tension in the AI spending chain: cloud providers continue to raise prices (AWS hiked its GPU reservation service by ~20% in July, following a 15% increase in January), yet corporate CFOs are beginning to scrutinize the ROI of every token burned.

Demand Has Not Vanished — It Is Becoming Selective
Despite Meta's signal to sell compute, top cloud vendors continue to report robust revenue and capital expenditure guidance. AWS Q1 revenue hit $37.6 billion (+28% YoY), Google Cloud reached $20 billion, and Microsoft Azure maintained ~40% growth. Amazon, Alphabet, and Meta each raised their full-year capex guidance to $200 billion, $180-190 billion, and $125-145 billion respectively. These numbers do not support a demand collapse story. Instead, they point to a bifurcation: cloud providers sell 'roads' — as long as models and enterprise customers run on them, they can collect tolls regardless of which model wins.

However, pure-play model companies face a different reality. Anthropic sees compute as a bottleneck because users trust it with expensive production tasks (coding, long-running workflows); interestingly, some compute from Elon Musk's xAI has even been redirected to Anthropic — machines care only about utilization, not founders. Meanwhile, Google reportedly restricted Meta from using Gemini because Meta's demand exceeded what Google could supply, showing that even Meta itself struggles to access top-tier model capacity. This is not classical overcapacity but a mismatch: high-certainty, high-scarcity compute remains in tight supply (AWS's two price hikes prove it), while mid-tier and low-tier compute face the brutal question of utilization.

UBS's recent survey of enterprise IT executives reveals that about 60% of companies are actively capping token spend and adding usage guardrails, especially after moving AI from pilot to production. The survey indicates that as AI shifts from a 'toy' to a 'tool', spending becomes harder rather than easier. CFOs no longer ask how many tokens were consumed; they ask how much money those tokens saved or earned.

Compute Has Not Disappeared — It Is Stratifying
A study by OpenAI and several universities on Codex provides staggering data: active Codex users grew more than 5x in the first half of 2026; median monthly output tokens for internal legal roles increased 13x compared to November 2025, and for research roles more than 50x. More importantly, agentic coding tasks can consume up to 1,000x more tokens than ordinary code chat or code reasoning, with runtime variability as high as 30x across different runs. Token has become an 'electricity meter' — software agents act like armies of mini workers that never take lunch but constantly burn tokens. A fast-spinning meter can indicate a busy factory or an inefficient one.

The 'electricity meter' logic, combined with CFO scrutiny, pushes enterprises to search for cheaper models. Chinese open-source models — GLM-5.2 by Zhipu (智谱), Kimi 2.7 by Moonshot, DeepSeek, Qwen — are thus becoming price-cutting tools on enterprise procurement desks. Marc Andreessen of a16z publicly stated that many AI practitioners now consider GLM-5.2 capable of matching or even surpassing top American public models on most tasks. Coinbase CEO Brian Armstrong provided concrete evidence: after switching the default AI model to GLM 5.2 and Kimi 2.7, combined with model routing, caching, and context compression, token usage continued to grow exponentially while total AI spending dropped by nearly half. This is the first time enterprises can unbundle model procurement: the hardest tasks go to the most expensive frontier models; routine summarization, customer service, information extraction, templated code, and internal knowledge base queries go to cheaper, local or open-source models.

Meta's consideration of selling compute, therefore, is not an isolated event. Together with Palantir's criticism of token pricing and Coinbase's embrace of open-source models, it tells a single story: the AI spending chain is being disassembled. Upstream players sell certainty (cloud providers can keep raising prices); midstream players sell outcomes (model companies are judged by ROI); downstream players compress unit costs (open-source models serve as price anchors). Compute is no longer a total-narrative story but a structural one. The top layer remains tight (frontier models, premium cloud services); the middle layer grows awkward (not scarce enough, buyers comparison-shop); the bottom layer is squeezed by open-source models and cost optimization. Assets don't care whether you believe in the future — they only care whether they are utilized today. For the past two years, everyone rushed to acquire machines; now the market asks: who will find the killer app first? And even if they find it, who will actually make money? The AI capex narrative has officially shifted from 'who buys fastest' to 'who runs full'.

