Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain credit

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain credit

N
News Editor
2026-07-16 04:54:26
Galaxy Digital has laid out a broad framework for what it calls the “inference capital markets,” arguing that AI inference is moving from a purely technical service into an asset class that can be priced, hedged, financed and traded. In a research report written by Galaxy Digital Vice President of Research Lucas Tcheyan and circulated in Chinese by TechFlow, the firm links several parallel developments into one market structure: the rise of GPU price indexes, planned GPU futures from Intercontinental Exchange and CME Group, tokenized claims on future AI inference, useful proof-of-work networks that subsidize inference production, and stablecoin-funded lending against GPU hardware. The report’s central claim is that inference has now overtaken training as the main driver of global GPU demand, while autonomous agents are emerging as a new class of machine-native buyers that can pay for model output programmatically. Galaxy argues that the market is still early and fragmented. It sees progress on the off-chain side, where Ornn, Silicon Data and Compute Desk are building reference pricing for compute, and where Kalshi, ICE and CME are already moving toward tradable GPU-linked products. On-chain, the report highlights Venice’s VVV and DIEM system for tokenized inference access, Pearl and Ambient’s different attempts to turn inference production into useful proof-of-work, and USD.AI’s stablecoin-based credit model for financing AI hardware. Even so, the report says the sector has not yet solved its hardest questions: whether real demand for verifiable, censorship-resistant inference will grow beyond a niche, how token value can be tied to actual product usage instead of emissions and speculation, and whether legal enforcement and collateral recovery in GPU-backed lending can hold up in a true stress cycle.
Galaxy DigitalAI inferenceGPU futuresVenicePearlAmbientUSD.AICompute finance

Galaxy Digital has published a lengthy research report arguing that a new market structure is forming around AI inference, one that treats compute not just as infrastructure but as something that can be owned, financed, priced and traded. The report, written by Galaxy Digital Vice President of Research Lucas Tcheyan and translated into Chinese by TechFlow, describes an “inference capital market” as an emerging system of networks, protocols, infrastructure and applications that move AI model inference away from centralized APIs controlled by frontier labs and hyperscalers, then coordinate settlement on-chain and build a financial layer on top.

In Galaxy’s framing, users can send prompts to networks of GPU operators that are coordinated and incentivized by crypto tokens. In some cases, those systems can also offer cryptographic or economic guarantees around correctness and privacy.

Inference, not training, is now the main driver of GPU demand

The report says interest in this category has accelerated in 2026 as inference, the process of applying trained AI models to new data and generating outputs, has overtaken training as the dominant share of global GPU demand. At the same time, autonomous agents are showing up as a new class of inference buyers. They can pay programmatically and operate without human intervention once deployed.

Galaxy argues that the pieces of this market already existed in isolation. Decentralized GPU marketplaces, inference protocols, payment rails, tokenization tools, capital formation mechanisms and on-chain liquidity each had their own moment over the past few years. The change now is convergence. Those primitives are being combined into one integrated system, and the report calls that system an inference capital market. If inference becomes a standard input for more and more work, Galaxy expects the demand side of that market to keep expanding, with activity not limited to crypto-native users.

The first force behind that convergence is the shift in GPU usage from training to inference. Open-weight models, in the report’s view, are catching up on “good enough” task tiers, making it possible to route jobs that were once expensive to the cheapest provider available, whether or not that provider is crypto-native. Growing demand for inference is also pushing buyers to search more aggressively for compute. Galaxy points to a recent Citadel report showing that token spending measured by the Silicon Data LLM Index is falling, which it reads as evidence that users are moving toward cheaper models. The report is explicit that these AI tokens are pricing units used by AI companies, not crypto tokens.

Galaxy also notes that Coinbase, Microsoft and AirBnB have recently shifted toward open-source models, especially Chinese models, and says OpenRouter’s recent fundraising points to rising demand for diversified model access. That demand matters because it can make inference cheaper. Part of the pressure is still on the supply side, since chip shortages continue to raise marginal inference costs.

The second force is financialization. As AI spreads and becomes a general-purpose intelligence input for a wide range of tasks, Galaxy says the market is starting to demand ways to commoditize and financialize compute. More teams are trying to turn AI capacity into a tradable asset and connect it to a broader financial layer. In that context, an early framework is beginning to appear for financializing AI hardware and capacity and assembling those pieces into something closer to a complete market.

GPU indexes and futures are developing first off-chain

Before getting to the on-chain stack, Galaxy spends considerable time on the off-chain market for GPU futures, which it sees as the larger foundation underneath everything else.

The report cites a wide range of estimates for AI infrastructure spending. Morgan Stanley forecasts roughly $2.9 trillion in global data-center capital expenditure by 2028, excluding power investment, with about $2.5 trillion tied to AI. McKinsey estimates that data centers will require $6.7 trillion in global capital expenditure by 2030, including $5.2 trillion for AI processing facilities and $1.5 trillion for traditional IT, with AI scenarios ranging from $3.7 trillion under constrained demand to $7.9 trillion under accelerated demand. Goldman Sachs estimates about $7.6 trillion in AI infrastructure capital expenditure from 2026 through 2031, spanning compute, data centers and power. The precise number varies by forecast, but the broad conclusion is the same: compute and hardware are the largest spending bucket, making up roughly 55% to 67%.

Galaxy says those forecasts are hard because the unknowns sit on both sides of the market. One is demand elasticity. If cheaper compute gets reinvested into larger models and broader deployments instead of simply lowering bills, efficiency gains could expand usage rather than reduce spending. Another is useful chip life. Depreciation estimates range from three years to seven years. New chips launch every year and should in theory make older chips obsolete, yet older hardware keeps retaining value because supply remains tight and lower-tier models can still run on it. In Galaxy’s view, that means large amounts of capital are flowing into a volatile asset class, exactly the setting where pricing, hedging and financing markets tend to develop.

The report quotes Baseten, one of the world’s largest inference operators, describing compute procurement as being like a drug market: “you have a guy,” and when you need supply, you call him. Galaxy’s point is that a market already exists, but not in standardized form. Large buyers are privately locking in future compute through hourly rentals, multi-year reserved contracts that resemble take-or-pay agreements for GPUs, and bilateral deals between providers and major customers. Pricing is often opaque and relationship-driven. Frontier labs sell tokens in bulk. Hyperscalers reserve capacity with each other. Neo-cloud providers buy forward from clouds and brokers because demand outstrips supply.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

Galaxy argues that firms that profit from opacity, especially brokers and large inventory holders, have little incentive to trade that system for visible screens and a few basis points of efficiency. It compares that resistance to the interests that helped stall attempts to build LNG exchanges over the last decade. GPU futures, in this reading, are emerging not as a replacement for the way capacity gets allocated, but as a standardized layer built on top of a fragmented market to transfer price risk.

For futures to work, the market needs a credible reference index. For compute, Galaxy says, that is much harder than for standardized commodities. A “GPU hour” means very little unless it specifies chip model, memory and network configuration, region, and whether capacity is on-demand or reserved. Power, bandwidth and LNG had similar underlying differences before they became liquid markets. The way those markets solved the problem was to define grades and reference prices rather than require every unit to be identical. Crude oil has WTI and Brent. Natural gas has Henry Hub.

Galaxy says GPU pricing is starting to move in that direction. Ornn, a Galaxy portfolio company, has published a compute price index based on real-time transaction data. Silicon Data publishes daily rental indexes for H100, A100 and B200 on Bloomberg terminals, normalizing pricing across configurations, vendors and regions into benchmark series. Compute Desk is building in the same direction. Using Ornn’s framing, Galaxy says these indexes resemble SOFR more than LIBOR because they are based on broad sets of actual market transactions rather than panel-based estimates, and because they track a defined basket of compute rather than a single unit.

Still, the report stresses that GPU benchmarks face a problem crude oil does not. A barrel of WTI does not change, but GPU references decay with every hardware generation. H100 becomes H200, then B200, GB200 and eventually Rubin, which means the benchmark itself needs to be rewritten from one generation to the next. Fragmentation makes the job harder. AMD chips, Google TPUs, Amazon Trainium, hyperscaler custom silicon and sovereign chips all spread demand across incompatible hardware stacks.

Galaxy flags settlement as another open debate. Labs that want to hedge compute budgets and trading desks that want directional price exposure may only need a cash-settled contract that pays the index difference. But neo-cloud providers that need actual chips to serve customers may need capacity, not just price exposure. So far, the contracts coming to market are cash-settled because price hedging is the easiest thing to standardize, much as it is in many commodity futures markets. Yet the report also presents the opposite argument: where a small number of sellers control supply, cash settlement against a thin index can be easy to manipulate, and commodities often need physical delivery or at least a workable cash-and-carry link so prices converge to underlying reality.

The market also needs participants with real business reasons to trade. Galaxy identifies natural buyers as AI labs, application companies and neo-cloud providers that have promised downstream capacity and need to secure inputs. Natural sellers are firms sitting on GPU inventories with uncertain future use, including hyperscalers, large GPU holders and brokers. Lenders financing GPU procurement also need a reference price because debt backed by depreciating hardware must be marked against something. Speculators and proprietary trading firms can add liquidity on top. For now, the report says, the biggest structural tension is that most sellers want to sell long-dated contracts to lock in revenue, while most buyers prefer shorter tenors to preserve flexibility.

Despite those challenges, Galaxy says early signs of a more mature GPU market are already visible. Prediction market platform Kalshi has launched markets tied to specific GPU prices. Intercontinental Exchange, the parent of the New York Stock Exchange, is working with Ornn, while CME is working with Silicon Data, and both have announced plans to launch GPU futures in the coming year. In Galaxy’s words, “compute as a commodity” is moving closer to reality.

On-chain inference capital markets are taking shape across several layers

Galaxy describes model and inference providers as token factories. They take raw inputs, GPUs, and refine them into tokenized outputs. GPU hours are moving toward standardization through indexes, but the token layer above them is much less mature because one model’s token pricing looks nothing like another’s. Even so, the report says this layer is beginning to form. China’s three major state-owned telecom carriers have already started retailing inference as a metered utility, selling standardized monthly token packages in a format that resembles mobile data plans. Amazon is reportedly preparing to pay Anthropic based on tokens consumed rather than previously committed compute hours. The Shanghai Futures Exchange is also reportedly in early design work on AI token futures, a possible mirror to the GPU input contracts that CME and ICE are building.

Crypto is building its own version of this stack. Galaxy says on-chain inference capital markets are being built on top of earlier crypto-AI primitives such as GPU suppliers and decentralized model developers, while also incorporating emerging verticals like agent payment standards and tokenized inference markets. The ecosystem spans multiple chains and execution environments, though the report says development is especially concentrated on Base and Solana because of their established developer and user bases.

At the center are inference providers and networks, the projects that turn prompts into outputs. Surrounding them are the layers that make inference useful, accessible and financeable: model developers, GPU and compute suppliers, routers and marketplaces, agents and applications, payment rails and capital formation infrastructure. Those surrounding layers matter because they either create demand for inference, supply the inputs needed to provide it, or transform usage into something that can be paid for, financed, routed or owned.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

Much of this is not uniquely crypto-native. Galaxy notes that many pieces have off-chain equivalents. Agent frameworks such as Hermes and Ironclaw can route between frontier lab APIs and on-chain providers like Venice. Models from decentralized developers including Nous Research can be accessed on OpenRouter. GPU suppliers act as permissionless, open versions of hyperscaler and data-center infrastructure, usually at much smaller scale. Agent payment protocols such as x402 and MPP can pay for OpenAI or Anthropic subscriptions just as easily as they can pay Venice. Programmatic settlement is becoming standard rather than a crypto-only advantage, and the report points out that OpenAI and Visa have both announced their own agent payment infrastructure.

What looks more distinctive to Galaxy is the financialization side. That is where crypto changes how inference gets owned, priced and financed. The report breaks those efforts into three categories:

  • Inference service providers such as Venice and Morpheus tokenize access to future inference and turn those claims into something that can be held, priced and resold.
  • Useful proof-of-work projects such as Pearl and Ambient tokenize inference production, paying token emissions to subsidize the cost of serving inference.
  • Credit providers such as USD.AI do something different. They do not tokenize inference itself. Instead, they finance the hardware needed to run inference, using stablecoin deposits to fund loans against GPUs and data-center infrastructure.

Taken together, Galaxy says, those components form the early on-chain inference capital market.

Inference providers: early traction appears, but scale is still limited

The inference-provider layer is the closest thing in crypto to the conventional AI API market. Users or developers choose a model, submit prompts, pay by token, per request or by subscription, and receive outputs. At the simplest level, that looks similar to OpenRouter, Together AI, Fireworks or a frontier-lab API. The difference is that crypto-native providers may source capacity from decentralized GPU networks, accept stablecoin or token payments, offer access to open or uncensored models, include privacy guarantees, or attach tokenized access rights to actual use.

Galaxy says OpenRouter is one of the best places to observe on-chain inference because demand there is priced by token and users can switch providers on any given request. That is the kind of market where cheaper or faster providers should take share. Over the past three months, on-chain providers handled about 0.5% to 1% of OpenRouter’s daily token volume, while OpenRouter’s overall token volume kept rising sharply. Galaxy reads that as evidence of some initial traction beyond the crypto-native community, but still a very small share of total usage. On that basis, the report says these providers are not yet competing head-on with mature centralized products, whether because of limited distribution, relative cost or other constraints.

OpenRouter captures only part of the picture, though. Galaxy cites Venice’s figure that on June 23 its access points processed 100 billion tokens in total, roughly 10 times the amount it processed through OpenRouter. Looking only at OpenRouter therefore understates project-level traction. The report says on-chain inference providers are trying different ways to build a stable customer base. Venice has leaned hard on privacy as a differentiator, pitching users on the idea that a provider should not retain, inspect, leak, censor or be compelled to disclose sensitive prompts. Chutes and AkashML let anyone connect GPUs to their networks and monetize idle capacity in an effort to lower cost. Even so, Galaxy cautions that many of these features can ultimately be copied by centralized providers and may not be enough on their own to win meaningful share.

Where on-chain products can create real differentiation, the report argues, is in mechanisms that financialize inference itself by turning access rights into assets that buyers can own, hold and resell rather than consume once through a subscription.

Venice and the attempt to make inference access ownable

Galaxy calls Venice the furthest-developed example of turning inference access into an ownable asset. Founded by Erik Voorhees, Venice runs a two-token system, VVV and DIEM, that packages claims on future inference into something holders can mint, own and resell.

VVV serves as the project’s “capital asset.” It is not a claim on Venice equity, which exists separately. The report notes that Venice raised a $65 million Series A in June at a unicorn valuation. Holders of VVV can in theory benefit from the project’s success in one direct way: part of Venice’s revenue is used to buy back and burn VVV. Those burns happen through two channels, discretionary burns funded from general revenue and programmatic burns that route a fixed share of each new subscription into buybacks. Galaxy says 42% of VVV has been burned so far.

VVV also has utility. Any amount can be staked to earn annual VVV emissions, and staking 100 VVV unlocks a Pro subscription. The more interesting function is its link to DIEM, which Venice describes as a compute asset. Holders lock staked VVV to mint DIEM, and each DIEM permanently grants $1 of Venice inference credits. Hold 100 DIEM and you have $100 of API credits across the Venice platform, permanently, or at least for as long as Venice is still operating.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

The amount of staked VVV required per DIEM follows a curve set by Venice and rises exponentially as DIEM supply approaches a target controlled by the platform. The report says that is because each DIEM is effectively a perpetual $1-per-day liability on Venice’s books. Since supply is now close to that target, the minting rate has risen from about 90 VVV per DIEM at launch to several hundred VVV today. That makes issuance harder and means early minters obtained DIEM at much lower VVV costs than are available now. While VVV is locked behind DIEM, stakers keep only 80% of normal VVV staking rewards, with the other 20% going to Venice. The lock can only be released by burning DIEM, so anyone who sold DIEM must buy it back on the market to redeem the underlying VVV, taking a loss if DIEM has risen in price.

Galaxy says the two tokens reinforce each other. DIEM can only be minted by locking staked VVV, so stronger DIEM demand removes VVV from circulating supply and gives it a use beyond pure speculation. In the other direction, DIEM benefits from Venice’s growth. The more useful and widely used the platform becomes, the more valuable a transferable claim on its daily inference access should be. Holders are not just sitting on resellable inference credits. They are holding a position linked to Venice’s success.

The report adds that the broader product can drive the token economy even when users never engage with crypto directly. According to the Venice team, most users are not crypto-native and many do not care about the tokens. But subscriptions, credit purchases and platform usage still drive VVV buybacks and demand for Venice inference. In Galaxy’s description, the token layer sits downstream of the product rather than replacing it.

What makes DIEM unusual is ownership. The report says it lets users own the inference they consume rather than rent it. That opens several use cases. Because the claim is tradable, holders with uneven demand can keep baseline access and sell or rent out days they do not need, recovering costs that would otherwise be lost under a pay-per-use model. Agents can hold DIEM directly, giving them a permissionless, ownable inference balance. The report mentions instant sales via Aerodrome and fixed-term rentals through marketplaces such as Surplus, UsePod, AntSeed and CarpeDiem.

Galaxy also relays a common Venice example: a user buys DIEM, uses it for one day of inference, then sells it the next day. If the price is flat, the inference was effectively free. If the price rose, the user even made money. The reverse is also true. If the price falls, the holder’s loss can be much larger than the cost of buying inference directly. That means some users are speculating on inference prices while consuming inference at the same time.

DIEM can also provide cost certainty. A business or agent with stable, predictable demand can use it to lock in compute cost in a way that resembles a multi-year cloud reservation contract. Galaxy uses the July 7 DIEM price of $1,270 as an example and says one DIEM at that price represented roughly four years of $1-per-day credits, meaning the buyer was prepaying about three and a half years of a perpetual cash flow. But the report immediately points to the contradiction. To gain that certainty, the buyer has to hold a volatile, dollar-denominated perpetual asset. By pricing a perpetual claim this way, DIEM implies a double-digit discount rate on Venice’s ability to keep serving inference over time, and the whole claim is only worth something if Venice remains capable of providing service for years.

Galaxy then lays out several weaknesses in the mechanism. Tokenized inference is most useful for issuers that need to pull demand forward and raise capital. Labs with the best models and real pricing power have little reason to tokenize because doing so would sacrifice price discrimination, reduce revenue captured from unused credits and limit repricing flexibility. DIEM also has no maturity date, no principal return and no collateral or reserve backing. Unlike a GPU-backed loan, it offers no contractual recourse if Venice fails to keep serving that $1 entitlement years from now. On top of that, DIEM is not a claim on a fixed quantity of inference. It is a claim on whatever Venice decides $1 of inference can buy. Venice sets token prices across models, and those prices can move with demand and availability. The risk is not only market direction. It also comes from platform discretion.

Galaxy raises an even deeper question: is this perpetual, dollar-denominated format actually the exposure that inference buyers want, or would they rather own a fixed-term claim priced in compute or tokens, or some hybrid of the two?

For now, the report says DIEM is still held mainly as a speculative asset rather than used primarily for inference access. Less than 50% of the weekly issued amount is used for inference. Venice’s own materials describe DIEM as a “range-bound perpetual asset” and classify buyers into API users, VVV holders extracting value without selling VVV, and speculators arbitraging spreads. The latter two categories account for the largest share of holders. Galaxy says the nearest centralized analogue is OpenAI’s Scale Tier, where users prepay for model throughput measured in token-per-minute terms for a fixed duration. But Scale Tier is account-bound and non-transferable. DIEM’s strength is the opposite: it can be held, resold and composed with the rest of the crypto inference stack. Galaxy suggests a better instrument might combine Scale Tier’s fixed term and compute denomination with DIEM’s ownership and transferability.

From Venice’s side, every circulating DIEM is a $1 unit of compute that the company must keep serving forever and cannot sell to anyone else. That is a liability, and the report says this is why Venice uses revenue to buy back tokens. VVV and DIEM are not trying to mirror Venice equity. They began as bootstrapping tools for user acquisition and now derive value from the compute claims they provide. In Galaxy’s telling, one side holds a claim and wants it to rise in value, while the other side carries an obligation and wants to manage it. That shared position in Venice compute, rather than any equity interest, is what aligns the system.

Pearl and Ambient: tokenizing inference production

Venice tokenizes access. Pearl and Ambient are trying to tokenize production. Galaxy groups them under useful proof-of-work, the idea that token emissions can subsidize inference by making the same work secure the chain and produce something customers actually want to buy.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

Pearl rewires mining around matrix multiplication

Pearl Network is a Layer 1 forked from the Bitcoin codebase. It keeps Bitcoin’s UTXO model and difficulty-adjustment mechanism but swaps SHA-256 hashing for matrix multiplication, the core operation behind AI training and inference. The pitch is simple: the same matrix multiplication that serves customer inference can also count as a mining attempt.

When an AI model answers a prompt, it is effectively multiplying large numeric grids underneath the surface. Pearl has miners take those exact grids, perturb them slightly by adding a random layer, and multiply the perturbed version. That heavy computation is submitted into the mining race. During the process, intermediate results are checked to see whether they fall below the difficulty target. If they do, the miner wins the block, under the same rule Bitcoin uses, except that the tested work here is real model-serving computation rather than a purposeless hash. Once the multiplication is done, a quick final step removes the random layer and leaves the exact inference result the customer wanted. In one action, the network produces both a real AI output and a chance to win a block reward.

Galaxy says two design choices make this plausible. Pearl ships as a plugin for vLLM, popular software that AI companies already use, so providers can enable it without rebuilding their systems. And because a winning entry must be exposed for network verification, Pearl wraps it in zero-knowledge proofs so customer prompts and proprietary model weights stay hidden. The extra overhead is described as modest. Pearl reports 0.5% to 10% additional work, and in launch testing with Llama-3.3-70B, its version ran as fast as, and in some configurations faster than, the standard version because of core engineering changes.

But the report says Pearl still has a major flaw. The protocol cannot distinguish useful computation, work that serves real inference requests, from useless computation, because matrix multiplication is valid whether or not any customer needs the result. Pearl’s white paper already assumed there would be miners doing exactly that just to farm block rewards. The launch appears to have confirmed the risk. Early mining enthusiasm pushed compute sharply higher with little visible sign of real-world inference demand being served.

There are, however, signs of traction. Galaxy points to Pearl’s May announcement with Together.ai, one of the leading inference and compute providers. The two launched an inference endpoint priced more than 25% below Together’s standard rate, with the discount funded by Pearl token rewards earned on the same hardware. Galaxy’s conclusion is that Pearl’s dual-purpose design only becomes genuinely productive when real, paid inference demand drives the compute side. Without that demand, block rewards simply attract speculative miners and the outcome looks like a different flavor of Bitcoin-style proof-of-work rather than useful output.

Ambient ties mining to paid tasks through auctions

Ambient takes the opposite design path. Rather than letting miners run arbitrary models, it standardizes the whole network around one large open-weight model and builds consensus around verifying the outputs of that model.

Where Pearl has miners competing by brute force on the same problem, Ambient has them competing through auctions. Users or agents post an inference task with a deadline and a price, effectively saying “do this within X minutes and I will pay Y,” and miners bid to take the job. The winning miner runs the query on the network model and posts a bond that is forfeited if delivery fails, so the system has an economic guarantee on quality and speed. A randomly selected group of validators then checks the result, with selection weighted by a history of useful work rather than by staked capital. Since miners are serving many separate tasks instead of all racing for one block, the network avoids the throughput bottleneck that classic proof-of-work creates. Ambient is built as a Solana fork that replaces stake with useful work and is designed to run at Solana-like speed.

The auction is also how Ambient expects to deliver competitive inference pricing. A conventional API provider has to recover the full cost of a request from the user payment. Ambient miners can get paid twice for the same unit of work: once by the user or agent that posted the job and again by the protocol’s reward for verified useful work. Because miners compete on tasks with explicit price and latency targets, Galaxy says they should bid at net cost after expected token rewards, not at gross cost before them. In practice, emissions subsidize the supply side and the auction mechanism should force most of that subsidy through to demand in the form of cheaper inference. The key difference from generic mining subsidies is that rewards are attached to jobs that someone actually posted and paid for. If the mechanism works, emissions are buying not just compute but cheaper, verified inference, which in turn can attract more usage, create more work for miners and support demand for the network token.

Galaxy says this is why Ambient claims to have solved Pearl’s unresolved issue. In Pearl, miners can earn block rewards from matrix multiplication whether customers wanted the output or not. In Ambient, miners only get tokens by winning tasks that somebody has actually posted and paid for, so mining and serving real inference become the same action by design.

Ambient also takes a distinctive approach to verification. The report frames the problem like this: if a miner says it ran your query on the agreed model, how do you know it did not secretly switch to a cheaper, lower-quality model to cut cost? Galaxy notes that this is not just a decentralized-market problem. Centralized providers have faced similar accusations. Ambient’s answer uses logits, the raw numerical scores a language model produces for every possible next token before choosing one. That stream of scores functions as a fingerprint of the exact model doing the computation, and it can be hashed into a short number for checking.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

To verify an output of thousands of tokens, a validator does not need to rerun the whole task. It can choose one random point in the text, ask the miner for the fingerprint at that point, and run the model for just one token there to see whether the fingerprint matches. One token of work checks thousands. Galaxy compares the logic to Bitcoin, where producing work is expensive but checking it is cheap. Ambient says this keeps verification overhead near 0.1%, versus zero-knowledge-proof approaches attempted elsewhere that can cost roughly 10 to 1,000 times more.

Useful proof-of-work still depends on real demand and token closure

Galaxy says these projects differ from other decentralized-compute efforts because the work that secures the chain is the same work customers want. If that mechanism holds, a single unit of energy buys both security and a sellable product. Mining becomes a second revenue stream on hardware that providers are already operating, and the outputs can be verified well enough that agents do not need to trust a provider completely in order to buy inference.

If real demand is not there, though, block rewards alone can attract miners, leaving proof-of-work networks full of compute that serves no customer. The work looks useful in form, but not in substance.

Galaxy splits the current challenge into two parts. The first is demand. Decentralized inference networks compete not only with centralized providers but with plain GPU rental markets, both of which are often cheaper, faster and free of crypto tokens. To win, these networks need buyers who specifically value minimized-trust inference: verifiable, censorship-resistant, neutral and not dependent on a provider that can walk away. The report says that slice of willingness to pay is still small today. It could expand quickly if the projects prove they can provide stable, consistent inference at lower cost, or if trust in centralized AI erodes, but the market is not there yet. Pearl’s launch serves as the cautionary example. Without enough real demand, block rewards alone can pull in miners and fill the network with compute that serves nobody.

The second issue is token value capture. Each project promises some version of the same flywheel: real usage drives demand for the token, the token funds mining rewards that secure the network, and that supports more usage. Galaxy says none of them has actually closed that loop yet. Mining mints tokens and miners sell them to cover costs, but nothing on the demand side forces buyers to acquire the token because consuming the actual product, whether inference or proofs, mostly does not require crypto at scale. Pearl inference can be paid for in dollars. Ambient has delayed publishing its token economics and has not said whether inference will be denominated in the token. So the tokens are being earned and sold, not broadly used.

The most likely direction, Galaxy says, is that these networks will eventually make their tokens the native payment rails for inference. That would be the clearest way to close the loop. Combined with emissions that allow them to price below market, the strategy could be compelling: cheaper inference attracts real usage, and if usage must be paid for in the token, that usage becomes token demand. But the flywheel only turns in a healthy direction if sustained, organic token demand eventually grows larger than emission-driven sell pressure.

USD.AI and the on-chain credit market for AI hardware

Below tokenized access and tokenized production, Galaxy identifies a third on-chain market: financing the GPUs required to run inference. This is where the report sees crypto doing what it does best, precisely because the model does not rely on minting another token to bootstrap demand. Instead, it raises capital in a more conventional way, uses hardware as collateral, channels stablecoin deposits into loans for operators buying GPUs, and repays depositors from lease cash flows.

The largest operators already finance fleets through bank credit lines, asset-backed securitizations and private credit. Galaxy points to CoreWeave’s multi-billion-dollar GPU-backed debt as the best-known example. Smaller neo-cloud providers have a harder time. They may own hardware and have contracted cash flows that could support debt, but they often lack the balance-sheet scale, treasury function and lender relationships needed to borrow quickly. USD.AI lends to them. Depositors fund the loans, lease income repays them, and interest flows back as depositor yield. Galaxy says the model has three features banks struggle to match: the lender side is open to anyone holding stablecoins instead of a closed credit fund, each loan becomes a composable on-chain instrument that can be staked, traded or used as collateral elsewhere, and collateral claims are represented on-chain even though actual enforcement still depends on traditional legal processes.

USD.AI runs on two tokens. Depositors mint USDai, a synthetic dollar backed by PayPal’s PYUSD, which is itself backed by U.S. Treasuries and cash. USDai does not pay yield and is meant to remain liquid and composable. To earn yield, depositors stake into sUSDai, whose value rises as rewards accrue. Yield comes from two sources: interest paid by GPU borrowers on active loans and Treasury income earned on idle reserves between deployments. With the loan book running at roughly half the size of reserves, Galaxy says staking yield is around 8%, and the protocol aims to reach 10% to 15% as more capital gets deployed.

The hard part of lending against physical GPUs is enforcement when a borrower defaults. The report says USD.AI once recorded each financed GPU as an ERC-721 NFT and described that NFT as a lawful document of title under Article 7 of the Uniform Commercial Code, with machines held by borrowers under a custodial arrangement and the NFT functioning as collateral. The framework was called CALIBER. The protocol later abandoned it after deciding it created too much commercial friction. Now the NFT represents the loan record. It carries service terms and routes repayments on-chain, but it does not itself transfer the legal claim on the collateral. Enforcement instead happens through ordinary off-chain loan documentation, while physical recovery still depends on the same operational stack any hardware lender would rely on: site inspections, installation proofs, collateral monitoring, lien filings and cooperation from data centers or custodians. Galaxy notes plainly that this enforcement path, and the broader lending model around it, has not been tested through a full bad-debt recovery cycle.

Galaxy maps the emerging market for AI inference as a financial asset, from GPU futures to tokenized access and on-chain

The report also points to the asset-liability mismatch between a liquid token and a three-year amortizing loan. Most real-world-asset credit protocols hide that mismatch by promising instant redemption and then break under stress. Galaxy cites the USD0++ depeg as an example. USD.AI does not promise instant exits. Redemptions are processed over 30-day windows based on principal that has already amortized, on a first-come, first-served basis. The protocol does not liquidate a performing loan just to fund withdrawals. On top of that sits a pricing queue modeled on Flashbots’ MEV-Boost design, allowing redeemers who want to skip the line to bid for priority, with those fees routed to holders who keep waiting. Loan terms resemble CMBS structures: 70% to 80% loan-to-value, borrower reserve accounts covering about three months of debt service, liquidation after two missed payments, and insured, monitored hardware recoverable through specialized partners.

Galaxy includes USD.AI in the report for another reason as well: it connects the credit layer to the pricing layer. Lenders financing GPUs need some standard for marking collateral. They need to know how fast hardware depreciates, what it can fetch in a forced sale, what advance rate is safe and how residual value can be hedged. Compute indexes and the futures curves now forming around them can provide that reference. In turn, lenders create real credit exposure that gives those prices a use beyond speculation. A GPU lender ultimately cares less about what spot rental prices are on a given day than about what a machine could be resold for if a loan goes bad. Liquid indexes and futures should eventually help answer that question.

USD.AI says about 95% of its loan book is backed by long-term take-or-pay contracts rather than spot rentals, so borrower debt-service capacity depends more on contracted counterparties than on daily rental prices. Two things stand between that risk and depositors today. The first is fast de-risking. With 70% to 80% LTV and a prefunded three-month debt-service reserve, effective LTV at origination falls into the low 60s. USD.AI says about one-quarter to one-third of its loans pay down in the first year, letting the protocol recover substantial exposure before hardware values drop too far. The second is insurance. Every new loan now includes impairment coverage from Barkr, whose AI-driven collateral valuation is backed and reinsured by Munich Re. If a loan defaults and the collateral sells below Barkr’s assessed value, the difference is paid to the protocol. Given the 80% LTV ceiling, USD.AI describes this as full coverage of outstanding debt.

Galaxy is careful not to overstate that protection. Insurance changes the risk distribution but does not erase risk. It shifts residual-value risk to a stronger counterparty, while adding fresh dependence on Barkr’s valuation model and on the continuing effectiveness of the reinsurance structure. The coverage and USD.AI’s off-chain enforcement process have not been tested in a real default wave. Loans still amortize over three years against a claimed seven-year useful life, and a faster hardware cycle would compress that cushion. The difference from a few months ago, Galaxy says, is that a bad liquidation now hits the insurer before it reaches depositors.

Galaxy’s bottom line: financing has found real demand first

In its conclusion, Galaxy says both the on-chain and off-chain inference capital markets remain small relative to the growth of the broader AI industry. For the on-chain products to scale, they have to show that the advantages they introduce are durable.

The report says those advantages are clear in principle. Tokenized access, as in Venice, turns claims on inference into bearer assets that can be held, resold, rented or assigned to agents instead of being tied to a revocable vendor account. Useful proof-of-work, as in Pearl and Ambient, uses token emissions to subsidize inference below market cost and make outputs verifiable so buyers do not need to trust providers not to switch models. Financing, as in USD.AI, turns illiquid GPU credit into a composable instrument that anyone with stablecoins can fund and exit more quickly than through the traditional credit system. Under all three sits a permissionless, programmatic stack, which Galaxy sees as especially well matched to agents, the class of consumers it expects could drive much of the future demand for on-chain capital-market inference. In the report’s view, crypto matters where ownership, neutrality, composability and access to capital matter.

The obstacles are just as obvious. Nobody has yet connected real demand for compute to real demand for crypto tokens. Production networks mint and sell tokens while using emissions to fund below-market inference. Tokenized access rights still trade more on speculation around the issuer than on usage, with DIEM held mainly as a bet on Venice rather than priced as consumed inference. Financing is the exception because it already serves real customers, neo-cloud providers that need capital and have cash flow to repay, so its returns come from actual demand being financed rather than tokens being minted to draw attention. Up to this point, Galaxy says, the financial layer has been more successful at attracting speculative capital than at creating self-sustaining, usage-driven demand.

That is why the report argues the real edge of on-chain inference capital markets is not direct competition with incumbent AI companies on the thing they already do best, serving inference at scale at low cost. The stronger case is in forming capital and reaching markets that traditional finance is too slow, too small or too unequipped to serve. Galaxy describes this as a recurring crypto pattern. Crypto rarely wins on the product, the exchange, the model or the application itself, but it repeatedly becomes the fastest way to build the financial layer around those things, whether through asset pricing, fractionalization, financing or settlement.

Inference, in Galaxy’s view, is the latest and largest version of that pattern. A multi-trillion-dollar asset class is being assembled in real time, while the market structure required to treat compute as a financial asset, indexes, futures, credit and tokenized capacity, barely exists. That absence is the opportunity. The financing layer works today, the report says, because it is the first part of the stack to find real demand. Everything else is still a wager that the same advantages will continue to move upward as compute itself becomes financialized.

The report ends on a simple point: the inference market may take years to mature, but the financial layer being built around it is forming now. A note added at the end says the section on USD.AI was updated on July 15 to reflect changes in how the protocol enforces claims on defaulted loans.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.