Pantera Capital partner Jay Yu said compute and data center infrastructure spending has already become a trillion-dollar track, comparable in scale to the U.S. consumer market and a major pillar of U.S. economic growth. As AI agents move deeper into the real economy, demand for compute keeps rising. Yet the actual process of sourcing GPUs remains highly traditional. Platforms such as SF Compute, Vast AI, and Runpod are already active, but a large share of real GPU-node transactions still takes place through community chat channels, OTC brokers, and customized bilateral agreements.

Yu said the financialization of GPU compute is still at an early stage. Like electricity, compute is not a fully homogeneous and interchangeable asset. It is constrained by three factors: SKU, time, and geography. Even so, over a five- to 10-year horizon, he argued that compute could evolve into a commodity asset similar to electricity or crude oil, becoming a global resource that underpins economic output in the AI era.
Electricity markets offer a template
The article argues that electricity is a useful reference point for forecasting how compute markets may develop. Both electricity and compute are heterogeneous, time-sensitive, and concentrated at specific nodes. Both also sit at the base of broader economic activity.
Beginning in the 1990s, the U.S. power system moved through a gradual privatization process. Vertically integrated utilities that once combined generation, transmission, and distribution were broken apart. Independent grid operators took over regional dispatch, and large power hubs such as PJM West, ERCOT North, and CAISO SP15 eventually became tradable commodity contracts on Intercontinental Exchange (ICE) and CME Group, serving as pricing benchmarks for the asset class.
At the physical level, the power market follows a three-layer structure: grid, operator, and node. At the top are physically separate regional grids. Each grid has a system operator, either a Regional Transmission Organization (RTO) or an Independent System Operator (ISO). These operators manage many local substation nodes. The article uses San Francisco as an example, noting that the city’s power supply is supported by multiple substations, including Embarcadero, Larkin, and Mission.
Prices at each node are determined through Locational Marginal Pricing, or LMP. The algorithm solves for dispatch under constraints that include generation schedules, end-user demand, transmission costs, and the physical laws governing power flows. In Yu’s framing, LMP is the link between physical delivery of one watt of electricity and tradable benchmarks such as PJM-West. Hub prices and index prices are effectively weighted averages of baskets of end-node LMP prices. The largest and most liquid benchmarks were eventually listed on venues such as CME and ICE and became reference prices for the broader market.
How that framework maps onto compute
Drawing on roughly 30 years of electricity-market development, Yu lays out several hypotheses for compute.
First, he expects compute markets to adopt a similar structure. In physical delivery, the closest equivalent to the power market’s grid-operator-node stack is hardware-service provider-cluster. If electric grids are separate physical systems, then compute markets may split first by chip type. H100, H200, B200, and B300 could each have their own benchmark family, even if those benchmarks remain linked.

At the operator layer, Yu points to compute service providers such as AWS, Nebius, Coreweave, SF Compute, and Ornn. These firms price different clusters differently, and those clusters vary by time, location, and SKU. That makes them the closest analogue to grid nodes. The mechanism most similar to LMP, he wrote, is a scheduling or routing algorithm that outputs dynamic instance pricing based on resource availability for a given hardware specification.
Second, he expects many benchmark indexes to appear over time, with the winners most likely to be the ones tied to the most liquid physical-delivery infrastructure. He argues that this matches what happened in power and oil, where the contracts that became category benchmarks were anchored by the most liquid points of physical delivery.
Third, Yu says compute markets may also settle into a structure where OTC activity remains central. In power markets, generators and load-serving entities sometimes hedge directly on exchanges, but much of the market runs through bank and dealer OTC desks. Those desks convert nodal delivery into hub delivery at a spread and then manage their own risk. He expects a similar arrangement in compute, where the real commercial participants, namely neocloud providers and AI companies, may prefer to source specific hardware through brokers rather than hedge directly on an exchange.
Fourth, he believes basis risk in compute could be materially higher than in electricity. Under the Federal Power Act, the Federal Energy Regulatory Commission, or FERC, requires disclosure of wholesale electricity market data. No similar mandatory transparency regime exists for compute. As a result, dynamic compute prices and compute indexes can only be built from proprietary order books, including those assembled through revenue-sharing relationships and data purchase agreements with neocloud providers.
The inference stack creates structural longs and shorts
The article says the underlying driver of compute markets is token demand, especially demand for inference tokens. Yu divides the current inference stack into three broad layers.
- The neocloud layer, where firms such as Nebius and Coreweave operate physical data centers and act as sellers of GPU capacity.
- The on-tap layer, where developer platforms such as Fireworks and Baseten wrap bare-metal hardware into GPU environments that developers can use directly, making this layer a buyer of GPU capacity.
- The application layer, where products such as Cursor, Perplexity, and Rime deliver services to consumer and enterprise users. This layer buys inference tokens, whose underlying input is still GPU compute.
In Yu’s description, neocloud providers are the structural shorts in the GPU market, while the on-tap and application layers are the structural longs. Hyperscalers such as Amazon and Google, by contrast, are building products across every layer of the stack.
He also includes an approximate view of how profit flows through that chain. For every $100 that an upper-layer application spends on tokens, about $45 flows to the on-tap layer, $50 goes to the neocloud or GPU resource layer, and the remaining $5 goes to routing layers such as OpenRouter.

A possible capital-market structure for compute
That underlying structure, the article argues, shapes how compute capital markets may form.
On the long side are developer platforms and application companies. On the short side are neocloud providers. These two groups may trade specific SKUs through compute brokers and OTC desks such as SF Compute, Runpod, and Compute Exchange. Those brokers and desks would then manage inventory and absorb the basis risk between the hardware specification demanded by token consumers and a standardized H200 compute contract. They could hedge that exposure on exchanges such as Architect and Pluto.
Pricing on those exchanges, according to Yu, would likely be derived from weighted benchmark prices built from the order books of partner neocloud providers and OTC venues.
Nvidia as the compute market’s “central bank”
The article also gives Nvidia a special role. Nvidia recently said it plans to turn AI factory compute into an “investable asset.” In Yu’s framework, Nvidia can be viewed as the “central bank” of compute because of its dominant position in the GPU stack.
He breaks that comparison into three functions.
- Managing “inflation”: in the GPU economy, inflation is compared with the depreciation cycle of chips, especially the rate at which they lose value relative to state-of-the-art hardware. Nvidia can influence long-term depreciation by controlling release cadence for new chips such as Vera Rubin.
- Pursuing “full employment”: Nvidia benefits from high GPU utilization. If token demand lifts utilization, that creates more demand for GPU purchases. In that sense, Nvidia has an incentive to push financialization because more liquidity in GPU hours can translate into higher hardware utilization.
- Acting as “lender of last resort”: Nvidia has introduced a residual-value support policy of as much as 25%. Yu says that if neocloud providers or other GPU service operators face liquidity stress, the policy can be seen as a form of support or insurance for GPU asset values.
Four product categories are taking shape
As the financial layer around compute expands, the article groups the market’s likely products into four categories.
1. Physical delivery
This means delivering actual GPU resources to end users. Companies named in this segment include SF Compute, Hyperbolic, Vast, Runpod, and Compute Exchange. It is the layer closest to the hardware and the one that connects most directly to the AI trading participants described earlier.

Because hardware specifications vary so widely, many platforms begin as brokers that match demand with neocloud providers and collect commissions. The long-term goal is a spot exchange for compute. But Yu says it is difficult to maintain stable quality control in physical GPU delivery, especially for platforms whose supply comes from decentralized networks. For this layer to mature, he suggests the market may need something like a Moody’s-style ratings agency to certify the quality of delivered resources.
Even with heavy competition, he says physical delivery is also the layer most likely to build a durable moat. In oil and power, indexes, exchanges, and lending products tended to grow around the most liquid physical-delivery venues.
2. Index products
This layer consists of index curves built on top of compute assets. Yu names Ornn, Silicon Data, Compute Desk, and Semianalysis as firms working in this area. He argues that index construction is a critical step in making compute tradable and more fully financialized. As these indexes mature, the “compute market” narrative has gained traction, and large exchanges such as CME and ICE have announced cooperation with compute-index providers.
Still, large gaps remain between many compute indexes and prices seen in independent GPU markets. Yu gives two reasons. First, the underlying products differ: reliability, interruptibility, and contract terms vary widely across platforms. Second, the order-book data feeding different indexes are not the same, and compute lacks a mandatory transparency regime like the one imposed on electricity markets under FERC oversight. Many compute indexes today are built from bulk order-book purchases from neocloud firms or from datasets assembled through partnerships and revenue-sharing arrangements.
He also says the index layer is harder to monetize than exchanges or physical delivery. An index-only business may end up squeezed from below by the physical market that supplies price data and from above by exchanges seeking a share of economics.
3. Derivatives venues
The third layer is the trading venue, mainly cash-settled futures, built on top of indexes to provide hedging tools for compute. Liquid Compute and Architect are among the projects developing products in this area.
Yu says derivatives may attract the most attention and show the strongest commercial potential over time, but the market remains early and volumes are still limited. He also notes that the tradable instrument on these venues is usually a standardized H200 compute instance, not a hardware specification that can directly run inference tasks. That means the most active users are likely to be OTC desks and compute brokers hedging inventory risk on their balance sheets, rather than end-user AI buyers seeking immediate access to specific GPU resources.

4. Financial wrappers and funding tools
The fourth category is broader and includes financing structures for compute, such as lending protocols, treasury vehicles, and synthetic stablecoins like USD.AI. These products aim to support the physical buildout of data centers and neocloud providers. Yu also says the market may eventually produce risk-transfer and insurance products designed to smooth basis risk in GPU markets.
Why full-stack operators may have an edge
Put together, these layers suggest a fairly clear picture of what a mature compute industry could look like. Yu argues that the eventual winners may be full-stack firms: companies that can deliver GPUs physically, generate compute indexes from their own spot order books, and then build futures venues on top of those benchmarks so compute becomes an asset that can be hedged.
One unresolved question is sequence. The article says the market still has not settled whether the winning strategy starts with physical delivery, with indexes, or with exchange infrastructure.
Compute trading is not new, and the value chain is shifting
The article then widens the frame from GPU markets to what it calls the broader token economics of compute.
Historically, compute trading is not a brand-new idea. During 2023 and 2024, projects such as SF Compute and Hyperbolic had already laid out similar concepts, while decentralized compute platforms such as Akash and IONet emerged over the same period. A notable point, Yu says, is that many participants did not stop at GPU trading alone. Hyperbolic shifted into inference services and started selling inference tokens directly. SF Compute began building its own clusters and operating compute resources itself.
Along the value chain, value flows from AI applications such as Cursor, Harvey, and Granola, to routing layers such as OpenRouter, to inference providers such as Fireworks and Baseten, to neocloud providers and compute trading markets including Nebius, Coreweave, and SF Compute, and then to data center operators. In his view, companies that remain stuck in the middle of this chain are exposed to pressure from both directions. To preserve stable and attractive margins, they may need to expand upstream or downstream.
Three variables could reset the pricing curve
From a game-theory perspective, Yu identifies three variables that cannot be ignored in token markets.

- Nvidia controls chip supply and the pace of depreciation.
- Frontier labs such as Anthropic and OpenAI can release new models that raise demand for inference tokens.
- Open-source models such as Kimi, GLM, DeepSeek, and Tongyi Qianwen continue to put pressure on inference margins.
He says these variables together form the structural long and short drivers on both sides of the GPU and token markets. A major launch by any one of them can reset the entire pricing curve. Large events, including a release such as GPT 5.6, could trigger chain reactions across the AI token value chain and lead to sharp moves in both token prices and physical GPU prices.
That is why, in his view, it makes strong economic sense for one company to operate both an inference-token business and a physical GPU trading market. Doing so can offset swings in margins across adjacent layers of the chain.
From infrastructure to a financial asset
The article closes by comparing compute with the early days of oil and electricity markets. Those markets also began with bilateral, opaque, intermediary-led OTC trading. Over time, they developed standardized delivery agreements, benchmark indexes, trading conventions, and quality-verification systems, and only then became fully financialized asset classes.
Yu says compute is now moving through the same process. Projects are already testing indexes, cash-settled exchanges, lending products, and insurance tools, while physical GPU trading platforms and cloud providers are trying to capture more of the token value chain.
His conclusion is that compute, as a foundational input for the AI economy, is starting to move from pure infrastructure toward an independent financial asset. He says the sector could produce multiple unicorns across physical supply, brokerage, lending, and risk management, with business models that can connect both decentralized finance and traditional capital markets.
The article ends with a broader claim: compute may be the first genuinely new physical commodity to emerge in decades, and the market is now watching it move from fragmented community trading toward a mature asset class.

