Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure

N
News Editor
2026-07-02 22:01:05
Meta's consideration of leasing out spare GPU capacity triggered a sharp selloff in AI stocks, reigniting fears of capital expenditure overcapacity. On CNBC, Palantir CEO Alex Karp launched a blistering critique of the token-based pricing model used by OpenAI and Anthropic, calling it a 'wealth tax' on enterprises that pay for tokens 'creating zero value.' Meanwhile, Zhipu GLM-5.2 was recognized by a16z's Marc Andreessen as one of the first Chinese models to match leading US models on most tasks, and Coinbase slashed its AI costs by nearly half after switching to open-source models. This article argues that AI compute is not in overall surplus but is undergoing structural stratification: top-tier compute (flagship models, premium cloud) remains scarce, mid-tier is becoming awkwardly substitutable, and bottom-tier is being squeezed by open-source alternatives. Enterprise CFOs are scrutinizing token consumption, breaking the AI spending chain into three layers: selling certainty (cloud providers), selling results (strong models), and squeezing unit costs (enterprise procurement). The narrative has shifted from 'how much is spent' to 'how efficiently assets are utilized,' reminiscent of the railroad and fiber-optic bubbles where the direction was right but the timing was wrong.
AI ComputeCapital ExpenditureMetaPalantirZhipu GLMToken EconomicsOpen Source ModelsCloud ProvidersStructural Mismatch

The news that Meta is considering selling access to its surplus AI computing capacity sent shockwaves through AI-related stocks, triggering a sharp downturn. In a different era—say, three years ago—such a move would have been unremarkable. Cloud computing has always been about slicing servers and reselling them; Amazon, Microsoft, and Google have done it for years. But when Meta, the last major non-cloud player to amass a giant GPU fleet, starts talking about renting out its machines, the market begins to question whether the trillion-dollar AI capex cycle has overshot.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 2

On CNBC, Palantir CEO Alex Karp spent nearly 20 minutes on the topic. Announcing a new partnership with Nvidia, he quickly pivoted to criticize the token-based pricing models of OpenAI and Anthropic. Karp claimed that CEOs complained privately of 'paying for tokens that create zero value while handing over their proprietary data.' He described the rising model bills as a 'wealth tax' on enterprises.

Demand Hasn't Disappeared—It's Just Becoming Selective

Capital expenditure has been the core narrative of the AI bull market. As long as companies kept promising bigger spending, all segments rallied together. Meta's announcement triggered a reflex: 'Capex is collapsing.' But public data doesn't support that conclusion. AWS revenue grew 28% in Q1 to $37.6 billion—one of its fastest quarters in recent years. Google Cloud surged even more to $20 billion. Microsoft Azure maintained ~40% growth. Amazon signaled its 2025 capex could reach $200 billion; Alphabet raised its 2026 guidance to $180–190 billion; Meta itself raised its full-year capex to $125–145 billion. These are not collapse-level numbers.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 3

What is happening is a divergence. Cloud providers sell a 'road': as long as traffic flows, they collect tolls regardless of which car (model) drives on it. AWS even raised prices for its GPU-reservation service by about 20% effective July (after a 15% hike in January). That is not a sign of weak demand—it is a sign of scarcity.

Model companies, however, are not all equal. Anthropic is valued not because it is cheap but because users trust it with the most expensive tasks (coding, system changes, long-context reasoning). Its token consumption far exceeds casual chat. Strong models face a shortage of compute; weak models face indifference. xAI's Grok has not yet forged a clear enterprise identity, yet Musk's ecosystem redirected some compute to Anthropic—a stark reminder that hardware does not care about founders, only about utilization.

The Google-Meta relationship further underscores structural mismatch rather than pure surplus. In June, reports emerged that Google restricted Meta's access to Gemini because Meta's demand exceeded what Google could supply, even affecting some internal AI projects at Meta. A company simultaneously contemplating renting out its own compute while being unable to buy enough frontier model capacity—this is not oversupply; it is misallocation.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 4

The core problem facing enterprises now is: once they acquire compute, how do they justify the token bill to the CFO? A UBS survey found that about 60% of enterprise IT executives are tightening token spending and adding guardrails, especially those that have moved AI beyond the pilot phase. AI has shifted from a toy (where everyone feared missing out) to a tool (where every dollar must be justified).

AI agents amplify this concern. A Codex study by OpenAI and several universities showed that Codex's active users grew 5x in H1 2026; OpenAI's internal legal team's median token output was 13x higher than November 2025, and research teams 50x higher. Another study found that agentic coding tasks can consume up to 1,000x more tokens than standard code chat, and token consumption for the same task can vary by 30x across runs. The bottom line: software is becoming a swarm of tiny workers that constantly read files, run commands, edit code, fail, retry—and eat tokens at every step.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 5

When tokens become a meter reading, those who own the power plant have pricing power, but those who waste power get audited. CFOs will inevitably ask: which tasks require the strongest model, and which can use a cheaper alternative? At this point, open-source models like Zhipu GLM-5.2, Kimi, DeepSeek, and Qwen cease to be mere tech news and become leverage in enterprise procurement. Marc Andreessen of a16z noted that many AI practitioners now regard GLM-5.2 as one of the first Chinese models that matches or surpasses leading US public models on most tasks. Coinbase CEO Brian Armstrong provided the sharpest proof: after switching its default model to GLM 5.2 and Kimi 2.7, combined with model routing, caching, and context compression, token usage continued to grow exponentially, yet AI costs were cut by nearly half. For the first time, enterprises can disaggregate model capabilities: the hardest tasks go to the most expensive models; routine summarization, customer service, information extraction, and internal Q&A go to cheap or locally deployed models.

Compute Hasn't Disappeared—It's Just Stratifying

Meta's compute-for-lease story is not isolated. Together with Palantir's critique of token pricing and Coinbase's open-source pivot, it signals that the AI spending chain is being unbundled. The top layer (cloud providers) sells certainty; the middle layer (strong model companies) sells results; the bottom layer (enterprise procurement) squeezes unit costs. Every segment is still growing, but each is now being asked: 'Is the money well spent?'

For the past two years, the easiest story has been 'not enough': not enough GPUs, not enough electricity, not enough data centers, not enough engineers. In a gold rush, nobody counts pennies. But Meta's move forces a different question: once you own the machines, they don't automatically become good businesses just because they were expensive. They need daily utilization, paying customers, models that fill them, and applications that convert cost into revenue. This is 'utilization rate'—a cold, unforgiving metric that cares only about whether the machine is running today, not about your roadmap or your keynote.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 6

Cloud providers answer this question easily: they sell infrastructure, and everything ends up on someone's cloud. Strong model companies also have an answer: if the model is good enough, users will queue, enterprises will integrate, and compute becomes a bottleneck rather than inventory. The hardest position is the middle layer: companies with big fleets, big budgets, and model teams, but whose models are not front-runner, whose products are not daily habits, and whose developers do not adapt their workflows. For them, compute can flip from weapon to inventory with a single failed model release or user migration. Inventory must be discounted, rented out, or repurposed.

Therefore, the correct framing is not 'compute oversupply' but 'compute stratification.' The top layer remains tight: the best models, best clouds, most stable GPU clusters are still fought over (AWS can raise prices because certainty itself is a premium). The middle layer becomes awkward: it's not bad, but not scarce enough; customers compare, negotiate, and ask why they shouldn't use cheaper models. The bottom layer is squeezed by open-source alternatives and cost optimization: enterprises will not pay premium token prices for routine tasks; they will route, cache, compress, and tier models.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 7

Demand has grown up. Children don't check the bill; adults do. As AI enters the enterprise, it goes through a similar maturation: pilot phase is about FOMO; scale phase is about ROI. Once the math starts, the industry will no longer move in lockstep. Some players continue raising prices (selling irreplaceable certainty). Some shift to outcome-based pricing (because customers don't want to pay for consumption itself). Some are forced to cut prices (because acceptable substitutes exist). Some rent out capacity (because idle machines hurt more than discounted rates).

These simultaneous developments create the apparent contradictions: compute is scarce and also for lease; token consumption is exploding yet enterprise AI spending is being capped; frontier models are getting stronger while open-source models are getting cheaper. They are not contradictions—they signal that AI has moved from a 'total volume' story to a 'structural' story.

During the 19th-century railroad bubble, railroads were real. During the dot-com fiber bubble, fiber was real. Many railroad and fiber investors still lost money—not because the direction was wrong, but because they built too early, too much, or borrowed too expensively to wait for demand that took decades to arrive. AI data centers may also leave behind useful assets: GPUs will depreciate, power contracts will renew, and software will learn to eat more compute. But assets have their own temperament: they don't care about your belief in the future; they only care whether someone shows up to use them every day.

Meta Sells Compute, Palantir Slams Tokens, Zhipu Rises in Silicon Valley: AI Capex Shifts from Totality to Structure 8

Meta's compute-lease signal is stuck right there. It is not the end of AI, nor the end of semiconductors. It is the moment when, midway through the capex story, someone opens the warehouse door and lets the outside see how many machines are stacked inside. Some will be consumed by frontier models. Some will be rented by cloud customers. Some will become cheap in price wars. Some will wait quietly for an application that hasn't been invented yet.

For two years, the market believed all machines would eventually find their destiny. Now it is asking: who will find it first? Who will wait too long? And who will find it but still fail to make enough profit from it? Once that question is asked, the AI story no longer belongs solely to those who buy machines fastest.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.