Meta's Compute Sale Triggers Panic, Palantir Slams Token Model
The AI market experienced another violent correction after Meta signaled it might sell its excess AI compute. Three years ago, this would have been just another cloud business move, but coming from Meta—a company that bought chips for its own models, ads, and recommendations—it smells like a balance sheet painkiller. On the same day, Palantir CEO Alex Karp went on CNBC and spent nearly 20 minutes attacking the token-based pricing model of OpenAI and Anthropic. He claimed that enterprise CEOs privately complain about paying for tokens that create no value while surrendering their data, calling the soaring bills a "wealth tax" on businesses.

These two events together mark a turning point in the AI CapEx narrative. For the past two years, everyone focused on who dared to spend the most and fastest. Now the question is whether those machines can be kept fully utilized after purchase.

Demand Isn't Vanishing—It's Becoming Selective
The immediate fear from Meta's announcement was that "AI CapEx is collapsing," but public data does not support this. AWS Q1 revenue rose 28% to $37.6 billion, Google Cloud hit $20 billion, and Microsoft Azure maintained ~40% growth. Amazon hinted 2025 CapEx could reach $200 billion, Alphabet lifted its 2026 guidance to $180-190 billion, and Meta itself raised its full-year CapEx to $125-145 billion. These numbers do not signal demand collapse; they signal a divergence.

Cloud vendors sell infrastructure certainty—no matter which model runs on top, they collect tolls. AWS even raised prices for its GPU reservation service by ~20% in late June (following a 15% hike in January), a sign of scarcity, not weakness. But model companies face a different reality: Anthropic has strong models but insufficient hardware, while weaker models find no buyers. Some compute from Musk's xAI system is being diverted to Anthropic, showing that machines care only about utilization, not founders.

Enterprises are waking up. A UBS survey found ~60% of enterprise IT executives are curbing token spending and adding usage guardrails, especially those past the trial phase. Meanwhile, a Codex study by OpenAI and universities showed active Codex users grew 5x in H1 2026; median monthly output tokens for legal roles surged 13x and for research roles over 50x versus November 2025. Agentic coding tasks can consume 1000x more tokens than standard code chat, with up to 30x variance across runs. Tokens have become an electricity meter, and CFOs are asking what each kWh produces.
In this environment, Chinese open-source models like Zhipu's GLM-5.2 are becoming bargaining chips for enterprises. a16z co-founder Marc Andreessen noted that many AI practitioners now consider GLM-5.2 competitive with or even surpassing top US open models on most tasks. Coinbase CEO Brian Armstrong reported that after switching its default model to GLM 5.2 and Kimi 2.7, combined with model routing and caching, token usage continues to grow exponentially while AI spending dropped by nearly half. Open-source models don't need to win every battle—they just need to convince procurement that not every kWh must be paid at luxury rates.

Compute Isn't Disappearing—It's Tiering
The best interpretation of Meta's compute sale is not "compute oversupply" but "compute tiering." The top tier remains tight: the strongest models and most stable GPU clusters are still in high demand, and AWS can raise prices because certainty itself has a price. The middle tier becomes awkward: adequate but not scarce, customers compare, negotiate, and ask why not use cheaper models or another cloud. The bottom tier is squeezed by open-source models and cost optimization—enterprises will route, cache, compress context, and split models into different price tiers.

Analogies to the 19th-century railroad bubble and the dot-com fiber bubble are apt: the infrastructure itself was not wrong—many valuable networks grew on those rails—but many who built them lost money because they built too early, too much, or with too much debt. AI data centers will leave behind depreciated GPUs and power contracts. Assets don't care about your vision; they care only about daily utilization.

Meta's compute sale signals the AI CapEx narrative moving from a volume story to a structure story. Some players will continue raising prices (selling irreplaceable certainty), others will shift to selling outcomes (because clients don't want to pay for consumption), some will be forced to cut prices (when adequate alternatives emerge), and some will rent out machines (because idle hardware is worse than low-margin rentals). These four dynamics happening simultaneously make the market look contradictory, but it is the natural maturation of a scaling phase. The market will no longer belong solely to those who buy machines fastest, but to those who can keep their machines generating sustainable cash flow.

