AI infrastructure is moving into a leverage phase
One of the biggest shifts in AI infrastructure over the past few years is that capital spending has started to outpace internal cash-flow growth. From 2020 to 2023, AI-related capex at the top five cloud providers was roughly 20% to 30% of operating cash flow. By 2025, that ratio had climbed to nearly 94%. In 2026, the five companies had already confirmed more than $700 billion in combined capex.
These companies still generate strong cash, but once infrastructure spending scales in the hundreds of billions of dollars for an extended period, debt capital naturally takes a larger role. Compute assets begin to follow their own financing logic.
GPUs, servers and data centers must be funded before revenue shows up, while Neocloud operators usually do not have the same level of entity credit as hyperscalers. Lenders are therefore looking more at the project itself: Is the future cash flow stable enough? Can it cover principal and interest? If there is a default, how much asset value can be recovered?
Take-or-pay contracts are what make this structure work. Even if a customer does not fully use the capacity it reserved, it still has to pay the agreed fee. For operators, that locks in part of future revenue in advance. For lenders, it turns highly uncertain GPU utilization into a contractual cash flow that is easier to forecast.
CoreWeave is the clearest example of that model, with more than 98% of its revenue coming from take-or-pay contracts. As the share of multi-year contracts from investment-grade customers rises, its financing depends more on counterparty credit and contract coverage. Neoclouds such as Nebius and IREN, which also have large long-term customer contracts, later adopted similar structures.
That changes the risk ranking in GPU credit. Lenders first ask whether contract cash flow can service the debt, and only then do they assess how much GPU value can be recovered in a default. The number of GPUs sets the size of the collateral base; long-term orders decide what financing terms those assets can receive.
Once cash flow is locked, asset price risk remains
Take-or-pay improves visibility on debt service cash flow, but the economic value of GPUs still changes over time. Years later, how much revenue this hardware can still produce and how much it is worth remain balance-sheet risks.
GPUs are both production equipment and technology products. They can keep generating rental income, yet they also sit inside an extremely fast upgrade cycle. Once a new chip generation arrives, the effects on older hardware eventually show up in rent, utilization, renewal pricing and secondary-market value.
That is why accounting depreciation and economic depreciation rarely move in lockstep. Accounting can amortize equipment cost over a fixed life, while the market is constantly reassessing how much competitive compute a GPU can still produce. For lenders, the key variables come down to the cash flow left over the remaining loan term and the value that can be recovered in a default. Market doubts have already emerged around depreciation policies for data-center assets at large cloud providers.
This is also why GPU credit can stay stable for a while even when rents fall. As long as the customer keeps performing, spot-price changes do not immediately hit current debt service. Risk tends to concentrate around renewal, refinancing and default resolution. At those points, market rent and residual value reset the safety cushion.
If rent and equipment value fall faster than the loan amortizes, LTV rises again and the cushion narrows. Long-term contracts can push risk into the future, but they cannot lock in an asset’s full lifetime economics.
The current gap in compute credit is clear: contracts lock in part of customer payments, while GPU market prices still float.
The market now relies mainly on multi-year capacity contracts to lock in prices early. In the early days, one-year H100 contracts traded at a noticeable discount to spot. As spot prices fell and forward contract prices recovered, that spread narrowed sharply. Supply improvements and contract changes both matter, but the deeper issue is simpler: without standardized forwards, futures or swaps, industry participants are forced to let long-form contracts handle both procurement and price management.
Long-term contracts can allocate risk between two counterparties. If the market needs a persistent public price and a way to transfer risk across institutions, it will need more standardized financial instruments. As more GPU assets are supported by debt capital, that need becomes more obvious.
What the compute market actually needs to trade
Once compute enters the capital structure, there is a clear side that absorbs price swings. Neoclouds hold rent-downside risk, AI companies hold procurement-cost upside risk, and creditors hold residual-value and refinancing risk. Derivatives usually emerge from exactly those risks that balance sheets cannot absorb on their own.
Three participants, three exposures
Compute price swings ultimately hit three balance sheets. Neoclouds lock in GPU, data-center and financing costs up front, but future rent still gets repriced. When long-term coverage is thin, falling rent quickly compresses EBITDA and DSCR, which gives them the strongest incentive to sell forward and lock in revenue early. AI labs and inference platforms sit on the other side. Higher GPU prices directly squeeze gross margin, and in tight markets price risk and capacity risk often arrive together, so they need to lock in both cost and availability. Lenders care most about the remaining cushion between the loan balance and the GPU’s economic value. Without public indexes and a forward curve, LTV, refinancing and covenants are hard to manage dynamically. As hedging requirements begin to enter loan terms, demand for compute derivatives moves from voluntary corporate risk management into the financing system itself.
Which risks can actually make it into derivatives
If you estimate the total addressable market for compute derivatives using global GPU shipments, data-center capex or total AI infrastructure spend, the result can easily be off by an order of magnitude. The risks that can truly enter derivatives are only the ones still exposed to market prices and not already absorbed somewhere else.
A large share of GPUs owned and used internally by hyperscalers keep their risk on the owner’s own balance sheet. Capacity already locked under multi-year fixed-price contracts has also had its price risk allocated in bilateral agreements. Both carry economic risk, but neither necessarily needs to trade in a market every day.
So there are several filters between global compute demand and the real derivatives TAM. That is why the future size of the compute derivatives market is likely to depend much more on the share of merchant compute than on a simple linear relationship with global GPU installed base.
For a long time, many frontier models ran inside a small number of large labs and hyperscalers, and compute procurement was just as internalized. As open-source model performance improves, more companies can deploy models themselves, or send workloads to third-party Neoclouds and inference platforms. Demand that was once trapped inside a few large balance sheets is gradually becoming public-market order flow. That is why open-source models matter for financialization: first, more compute starts to produce real traded prices, which makes it possible to build an index and then a forward curve; second, demand shocks reach the spot market faster. If a popular model launches and many firms deploy it at once, the new demand hits third-party GPU capacity directly, and both availability and pricing can change quickly. The more marginal demand the public market absorbs, the more rental volatility becomes a real P&L risk that companies need to manage.
After DeepSeek V4 launched, H100 rental rates rose about 7.5% within two weeks. Around the launches of Kimi K3 and GLM 5.2, H100 and H200 rents also strengthened in a similar way. Popular open-source models can trigger a burst of deployment and inference demand in a short window. When that incremental load is mostly absorbed by Neoclouds and third-party inference platforms, the market usually sees availability tighten first, along with queues and quota restrictions, before that pressure shows up in on-demand rents.
So the real importance of open-source models is that they may raise the share of market-based procurement. Compute only becomes an indexable, forwardable and financially tradable price after enough external transactions pass through it. The variables that will ultimately determine the size of the compute derivatives market are: how much compute enters the public market, how much of that price remains floating, and how much of the risk can no longer stay on the original balance sheet.
Without storage arbitrage, future compute depends more on expectations and order flow
The hard part of compute forwards is simple: the market still does not have a trustworthy term structure. GPU-hours cannot be stored. An idle hour of H100 today disappears forever once the window passes, and you cannot buy it cheaply in spot and deliver it six months later. The spot-to-forward discipline that comes from inventory arbitrage in traditional commodities is much weaker here by nature.
That means the price of a GPU six months from now will depend far more on chip deliveries, data-center power availability, model efficiency and actual order strength at that time. Delayed hardware shipments can tighten future capacity quickly, while software optimization can release large amounts of effective compute without adding more GPUs. A compute forward curve therefore embeds the market’s view on hardware supply, energy constraints and algorithmic progress all at once.
The oil market has a much more stable storage-arbitrage framework. When the forward price is clearly above the spot price plus financing and storage costs, traders can buy spot, hold inventory and sell forward. That keeps the spread from drifting far away from carrying costs. Compute lacks that cash-and-carry constraint, so spot prices bind the forward curve less tightly.
Future prices will therefore be highly sensitive to marginal information. Delays in Blackwell shipments or in datacenter power delivery reduce expected supply. Model distillation, inference optimization or a new chip generation can raise effective compute across the network. Many of those shifts can rewrite future supply and demand within months.
That makes real order flow more valuable. Public spot quotes only tell the market what happened today. Long-term capacity orders expose future supply and demand earlier. A broker or dealer that constantly matches Neoclouds with AI labs can see which data centers are starting to have spare capacity and which buyers are willing to pay a premium for GPUs six months out. That kind of information eventually feeds directly into forward pricing.
The value of a compute index is therefore to compress highly fragmented, non-standard quotes into a pricing benchmark that contracts, credit and derivatives can all reference. If an index can keep appearing in long-term capacity contracts, loan valuations, OTC settlement and futures delivery, it will gradually become the market’s shared anchor. Whether index providers and dealers can build a moat later will depend largely on how close they are to real trades and forward orders.
GPU-hours and AI tokens will not financialize at the same pace
There are really two price layers in the AI stack. Upstream sells GPU capacity. Midstream inference platforms turn GPU-hours into model calls, then sell AI tokens, APIs or specific workloads downstream.
For inference platforms, profit depends on both sides at once. Higher GPU rents lift input costs, while lower AI token prices squeeze revenue. If both sides become tradable, inference platforms could theoretically lock in part of a compute-to-intelligence spread, much like a refiner manages crack spread in energy markets.
But GPU-hours and AI tokens are still far apart in terms of standardization. GPUs do have differences in model, region, interconnect, cluster size and SLA, but they can still be standardized step by step around a chip generation, delivery window and region. A single H100 or H200 still has relatively clear technical boundaries.
Tokens have far less stable economic meaning. Pricing purely by token count makes it hard to place different models and service levels into the same financial contract. One million tokens can come from completely different model capabilities, and may also imply very different latency, throughput and reliability. Large enterprise buyers often pay less than public API list prices, which makes posted prices a weak settlement benchmark on their own.
Technology deflation makes term pricing for AI tokens even harder. Model architecture, inference stacks and hardware efficiency keep improving, so the compute needed for the same quality task can drop noticeably in a short time. Once the horizon extends, the market is not only forecasting token supply and demand; it is also trying to estimate how many tokens and how much GPU time will be needed for one unit of future “intelligence.”
That is why standardization on the demand side may not begin with token futures. More likely, the first products will be model baskets, routers and workload pricing: a coding task, search task or similar job is defined as a relatively stable service unit, and the underlying system dynamically chooses the model and provider. Only after those workloads accumulate enough real transactions will the market be able to form indexes and derivatives.
GPU-hours are closer to a basic commodity. Workloads are closer to final economic demand. The standardization process between them will determine what the compute financial market ultimately trades.
This article is for informational purposes only and does not constitute investment advice. Readers should consider whether any opinion, view or conclusion here fits their own situation before making any investment decision. Investing involves risk, and you are solely responsible for your own decisions.

