A financial layer for AI inference is beginning to form
Galaxy Digital research vice president Lucas Tcheyan argues that an “on-chain inference capital market” is starting to emerge. In the long-form report, translated by TechFlow and republished by WuBlockchain, he describes it as a system of networks, protocols, infrastructure and applications that moves AI model inference out of centralized APIs controlled by frontier labs such as OpenAI and Anthropic and hyperscale cloud providers, then coordinates settlement on-chain and builds a financial layer on top.

In Tcheyan’s framing, users can send prompts to networks of GPU operators coordinated and incentivized by crypto tokens. In some setups, the output can also come with cryptographic or economic guarantees around correctness and privacy. He says the category drew much more attention in 2026 for two reasons: inference has overtaken training as the dominant share of global GPU demand, and autonomous agents have emerged as a new class of inference consumers that can pay programmatically without human intervention during runtime.
Over the last few years, decentralized GPU markets, inference protocols, payment rails, tokenization tools, capital formation mechanisms and on-chain liquidity each had their own moment. The change now, he writes, is that these pieces are no longer developing in isolation. They are being assembled into a more integrated market structure around inference.
The first force behind that convergence is the shift in how compute is used. GPU demand is moving decisively from training toward inference, while open-weight models are catching up with frontier systems on tasks where “good enough” performance is enough. That makes it easier to route work to the cheapest provider, whether the service is crypto-native or not.
Tcheyan points to a recent Citadel report showing that token spending measured by the Silicon Data LLM index has been falling, a sign that users are shifting toward cheaper models. He makes clear that AI tokens here are units used by AI companies to price services, not crypto tokens. He also notes that Coinbase, Microsoft and AirBnB have recently moved toward open-source models, mainly Chinese ones, and says OpenRouter’s recent funding points to rising demand for diversified access to models.
The second force is financialization. As AI becomes a general-purpose input across more tasks, teams are looking for ways to turn AI hardware and capacity into tradable assets that can plug into broader financial rails. In Tcheyan’s view, the early framework for an inference capital market is already becoming visible.
GPU indexes and futures are laying groundwork off-chain
Before getting to on-chain markets, the report spends significant time on GPU futures. Tcheyan cites three large estimates for AI infrastructure spending. Morgan Stanley projects roughly $2.9 trillion in global data-center capital expenditure by 2028, excluding power investment, with about $2.5 trillion tied to AI. McKinsey estimates $6.7 trillion in global data-center capex by 2030, including $5.2 trillion for AI processing facilities and $1.5 trillion for traditional IT, with AI scenarios ranging from $3.7 trillion under constrained demand to $7.9 trillion under accelerated demand. Goldman Sachs estimates about $7.6 trillion in AI infrastructure capital expenditure from 2026 through 2031, spanning compute, data centers and power.
The exact figures differ, but the takeaway is the same. Compute and hardware are the biggest spending category, making up roughly 55% to 67% of the total. Tcheyan says forecasts remain difficult because both demand and supply are full of unknowns. If cheaper compute gets reinvested into larger models and wider deployment rather than booked as savings, efficiency gains could expand usage instead of cutting bills. At the same time, estimates for effective chip lifespan range from three years to seven years.
He quotes Baseten describing compute procurement today as something like a drug market: “you have a guy,” and when you need supply, you call that person. That line is used to show that these markets already exist, just not in standardized form. Large buyers privately lock up future compute through hourly rentals, multi-year reserved-capacity contracts and bilateral deals between suppliers and major customers, usually with opaque pricing shaped by relationships. Frontier labs such as OpenAI sell tokens in bulk, hyperscalers reserve capacity with one another, and neocloud providers buy forward from clouds and brokers because supply is tight.
GPU futures, in his view, are emerging as a standardized layer on top of that fragmented base. Their role is to transfer price risk, not yet to replace the way capacity itself is allocated. For that to work, the market needs a credible index. A “GPU hour” means little unless it specifies chip type, memory and networking configuration, region, and whether the capacity is on-demand or reserved. Tcheyan compares the problem to the early development of oil and gas benchmarks, where the answer was not perfect sameness but accepted grades and reference prices.

He highlights several companies trying to build those benchmarks. Ornn, a Galaxy portfolio company, has published compute price indexes based on real-time transaction data. Silicon Data publishes daily rental indexes for H100, A100 and B200 chips on the Bloomberg terminal, normalizing prices across configurations, providers and regions into a single benchmark. Compute Desk is moving in the same direction. Using Ornn’s framing, Tcheyan says these indexes look more like SOFR than LIBOR because they are built on broad sets of actual transactions rather than panel estimates.
Still, GPU benchmarks face a problem oil does not. A barrel of WTI does not change, while GPU benchmarks decay and have to be rewritten as the market moves from H100 to H200, B200, GB200 and eventually Rubin. Fragmentation makes that harder. AMD chips, Google TPUs, Amazon Trainium, hyperscaler in-house silicon and sovereign chips all split demand across hardware that is not fully compatible.
The report also spends time on settlement design. A lab hedging its compute budget or a trading desk taking directional exposure may want only a cash-settled contract tied to an index. A neocloud provider that needs actual chips to serve customers cares about physical capacity. The early products announced so far lean toward cash settlement because pure price hedging is easier to standardize, though Tcheyan notes another view: where supply is concentrated and indexes are thin, cash settlement can be easier to manipulate, so physical delivery or a functioning cash-and-carry mechanism may be needed for prices to converge to reality.
Natural buyers in this market would include AI labs, application companies and neocloud providers that have already promised capacity downstream. Natural sellers include hyperscalers, large GPU holders and brokers with inventory but uncertain future use. Lenders also need reference prices, since debt backed by depreciating hardware has to be marked somehow. One structural tension, he writes, is already clear: sellers want long-dated contracts to lock in revenue, while buyers want shorter-dated contracts to preserve flexibility.
Despite the complications, Tcheyan says signs of a more mature GPU market are already visible. Kalshi has launched markets tied to specific GPU prices. Intercontinental Exchange, the parent of the New York Stock Exchange, has partnered with Ornn, while CME has partnered with Silicon Data, and both have said they plan to launch GPU futures within the next year.
The stack that makes up an on-chain inference capital market
Tcheyan then turns to crypto-native infrastructure. He describes models and inference providers as “token factories”: they take raw input, GPU time, and refine it into outputs priced in tokens. GPU-hour pricing is becoming more standardized through indexes, but the token layer on top remains highly fragmented, with very different pricing across models. Even so, he says that layer is beginning to take shape as well.
He notes that China’s three major state-owned telecom operators have started retailing inference as a metered utility, selling standardized monthly token packages in a way that resembles mobile data plans. Amazon is reportedly moving toward paying Anthropic based on tokens consumed rather than on previously committed compute hours. The Shanghai Futures Exchange is also said to be in the early stages of designing AI token futures, which Tcheyan presents as a counterpart to the GPU contracts being built by CME and ICE on the input side.
Crypto’s version of this market is being built on top of earlier crypto-AI primitives such as GPU suppliers and decentralized model developers, while adding newer verticals including agent payment standards and tokenized inference marketplaces. The ecosystem now spans several chains and execution environments, but Tcheyan says development is especially concentrated on Base and Solana because both have deeper developer and user bases.
At the center are inference providers and networks that turn prompts into outputs. Around them sit model developers, GPU and compute suppliers, routing and marketplace layers, agents and applications, payment channels and capital formation infrastructure. Those surrounding layers matter because they either create inference demand, supply inference inputs, or convert inference usage into something that can be paid for, financed, routed or owned.

He is careful to note that many of these products are not unique to crypto. Agent frameworks such as Hermes and Ironclaw can source inference either from frontier labs or from on-chain providers such as Venice. Models from decentralized developers like Nous Research can be accessed through OpenRouter. Agent payment protocols such as x402 and MPP can just as easily pay for OpenAI or Anthropic subscriptions as they can for Venice. Programmatic settlement is becoming an industry standard rather than a crypto-only edge, and Tcheyan points to recent agent payment infrastructure announcements from OpenAI and Visa.
The more distinctive part appears on the financialization side. He groups those projects into three buckets. First, inference providers such as Venice and Morpheus tokenize access to future inference, turning it into a claim that can be held, priced and resold. Second, useful proof-of-work projects such as Pearl and Ambient tokenize inference production itself, using token rewards to subsidize the cost of serving inference. Third, credit providers such as USD.AI do not tokenize inference, but finance the hardware needed to run it, using stablecoin deposits to fund GPUs and data centers.
Inference providers: some traction, but still a small share
In the provider layer, decentralized inference most closely resembles the existing AI API market. Users and developers choose a model, send prompts, pay per token, per request or through subscriptions, and receive outputs. The difference is that crypto-native providers may source capacity from decentralized GPU networks, accept stablecoins or tokens, give access to open or uncensored models, include privacy guarantees, or bind tokenized access rights to usage.
Tcheyan calls OpenRouter one of the clearest venues for observing this competition. Demand there is priced per token, and users can switch providers freely on any request, which should favor whichever service is cheaper or faster. Over the last three months, on-chain providers handled 0.5% to 1% of OpenRouter’s daily token volume, while OpenRouter’s aggregate token volume kept climbing quickly. He sees that as evidence of some traction beyond the crypto-native audience, though it remains a small slice of total usage and suggests these providers still have a long way to go against established centralized offerings.
OpenRouter, however, is only part of the picture. Venice reported that across all of its access points it processed 100 billion tokens on June 23, roughly 10 times its OpenRouter volume for that period. Tcheyan uses that example to argue that marketplace data alone can understate project-level traction. Providers are also trying different ways to build a customer base. Venice pushes privacy as a feature, emphasizing that users do not have to worry about the provider retaining, inspecting, leaking, censoring or being compelled to disclose sensitive inputs. Chutes and AkashML try to lower costs by letting anyone connect GPUs to their networks and monetize idle capacity.
Even so, the report argues that many of those features could be copied by centralized providers. The place where on-chain products may build a more defensible difference is in turning access rights themselves into financial assets.
Venice and tokenized ownership of inference access
Tcheyan treats Venice as the furthest-developed example of that idea. Founded by Erik Voorhees, Venice uses a two-token structure, VVV and DIEM, to package claims on future inference into assets that can be minted, held and resold. He stresses that VVV is not equity in Venice. The platform has a separate equity structure, and in June Venice completed a $65 million Series A round at what the report describes as a unicorn valuation.
VVV functions more like a “capital asset” for the project. Part of Venice’s revenue is used to buy back and burn VVV. That happens in two ways: discretionary burns funded from general revenue, and programmatic burns that route a fixed share of each new subscription into buybacks. According to the report, 42% of VVV has been burned so far.
VVV also has utility. Any amount can be staked to receive annual VVV emissions, and staking 100 VVV unlocks a Pro subscription. The more unusual use case is its relationship with DIEM, which the report calls Venice’s “compute asset.” Holders can lock staked VVV to mint DIEM, and each DIEM grants $1 of Venice inference credits per day in perpetuity. A holder with 100 DIEM therefore has access to $100 of daily API credits across all models on the platform, permanently, at least so long as Venice keeps operating.
The amount of staked VVV required to mint one DIEM follows a curve set by Venice, and rises exponentially as DIEM supply approaches the target controlled by the company. Tcheyan explains the reason simply: every DIEM is a perpetual $1-per-day liability on Venice’s books. Supply is now near that target, so the mint rate has risen from around 90 VVV per DIEM at launch to several hundred today. During the period when VVV is locked to back DIEM, stakers keep only 80% of their normal VVV staking rewards and the remaining 20% goes to Venice. Unlocking the VVV requires burning DIEM, which means anyone who minted DIEM and later sold it has to buy DIEM back in the market to redeem the original VVV position.

That design links the two tokens tightly. DIEM can only be created by locking staked VVV, so stronger DIEM demand removes VVV from circulation and gives it a use beyond speculation. DIEM, in turn, benefits if Venice grows more useful and more widely used, because the transferable claim on daily access becomes more valuable as the platform’s service becomes more attractive.
Tcheyan notes that Venice says most of its users are not crypto-native and many do not care much about tokens. Still, when they subscribe, buy credits or use the platform, those actions drive VVV buybacks and create demand for Venice inference. In his telling, the token economy sits downstream of the product rather than replacing the product.
The key distinction of DIEM is ownership. In a normal pay-as-you-go model, the buyer consumes inference and is left with nothing. A holder of tokenized access keeps an asset that can be retained, transferred or sold. Tcheyan lists several implications. A holder with uneven demand can keep baseline access and rent or sell excess days. Agents can hold DIEM directly and treat it as a permissionless, ownable inference balance. The position can be sold instantly through Aerodrome or rented for fixed terms through marketplaces such as Surplus, UsePod, AntSeed and CarpeDiem.
The Venice team often gives a more speculative illustration: buy DIEM, use one day of inference, then sell it the next day. If the price is unchanged, the inference was effectively free. If the price rose, the user even made money. Tcheyan notes the reverse case as well. If the price falls, the holder can lose far more than the cost of simply buying the inference directly.
DIEM can also offer a form of cost certainty. A company or agent with stable, predictable demand can lock in future inference spending in a way that resembles long-term reserved cloud contracts. Using the July 7 DIEM price of $1,270, Tcheyan estimates that one DIEM represented roughly four years of daily $1 credits, meaning the buyer was prepaying for about three and a half years of perpetual cash flow.
He then turns to the limitations. Tokenized inference is most useful for issuers that need to pull demand forward and raise capital. Labs with the best models and real pricing power have less reason to tokenize because doing so would reduce their flexibility around price discrimination, breakage and repricing. DIEM also has no maturity date that returns principal and no collateral or reserve backing, unlike GPU-backed lending. Holders are making an open-ended bet that Venice will still be delivering that daily $1 years from now. On top of that, DIEM is not a claim on a fixed amount of inference. It is a claim on whatever amount of inference Venice decides $1 should buy, since the platform sets token pricing across models and that pricing can move with demand and availability.
He pushes the question one level deeper: is a perpetual, dollar-denominated structure like DIEM actually the exposure inference buyers want, or would buyers rather hold something time-limited, denominated in compute or model tokens, or some hybrid of the two?
For now, DIEM appears to be held mainly as a speculative asset rather than used as an access tool. The report says weekly inference usage amounts to less than 50% of issued DIEM. Venice’s own materials describe DIEM as a “range-bound perpetual asset” and divide buyers into API users, VVV holders extracting value without selling VVV, and spread-arbitrage speculators, with the latter two groups representing the larger share of holders.
Tcheyan compares DIEM to OpenAI’s Scale Tier, a prepaid commitment for model throughput measured in tokens per minute and sold for a fixed term. Scale Tier is not ownable inference, though; it is account-bound and non-transferable inside OpenAI’s platform. DIEM goes in the opposite direction: it can be held, resold and combined with the rest of the crypto inference stack. His view is that a better tool may eventually combine Scale Tier’s term structure and compute denomination with DIEM’s ownership and transferability.

For Venice itself, every DIEM in circulation is a unit of compute it must continue to serve and can no longer sell to someone else. That makes DIEM a liability, which is one reason the company buys back tokens using revenue. Tcheyan’s broader point is that VVV and DIEM are not really trying to imitate equity. Their value comes from turning Venice compute into a shared position: holders want the claim to become more valuable, while Venice carries the obligation and has an incentive to manage it carefully.
Useful proof of work: Pearl and Ambient take opposite paths
If Venice tokenizes access rights, Pearl and Ambient tokenize inference production. Tcheyan groups both under useful proof of work, where the work that secures a chain is also work that customers want to buy. In principle, that means a unit of energy pays for both security and a marketable product.
Pearl Network is a Layer 1 forked from the Bitcoin codebase. It keeps Bitcoin’s UTXO model and difficulty adjustment but replaces SHA-256 hashing with matrix multiplication, the core operation behind AI inference and training. When a model answers a prompt, the underlying computation is matrix multiplication. Pearl has miners take those matrices, perturb them with a random layer, multiply the perturbed version, and continuously test intermediate results against the network’s difficulty target. If the result falls below the target, the miner wins the block. Once the multiplication is complete, the random layer is removed, leaving the exact inference result the customer wanted.
Two design choices are meant to make that viable. Pearl ships as a plugin for vLLM, a widely used inference engine, so providers can turn it on without rebuilding their systems. Because winning entries must be exposed for network verification, the process is wrapped in zero-knowledge proofs so prompts and proprietary model weights remain hidden. Pearl says this adds only 0.5% to 10% more work, and in launch testing on Llama-3.3-70B, the Pearl version ran as fast as the standard one and in some configurations even faster.
The problem, Tcheyan writes, is that the protocol cannot tell useful and useless work apart. A matrix multiplication counts even if no real customer wants the result. Pearl’s own white paper assumes a class of miners that will run useless computations purely to collect block rewards. The launch data confirmed that risk: early mining interest pushed compute higher very quickly, with little sign that much of it was serving real inference demand.
There are, however, signs of real-world traction. In May, Pearl announced a partnership with Together.ai, one of the larger inference and compute providers, to launch an endpoint priced more than 25% below Together’s standard rates. The discount is funded by Pearl token rewards earned on the same work. Tcheyan’s conclusion is straightforward: the design only produces genuinely useful work if paid inference demand is what drives compute in the first place. Without that, block rewards simply attract speculative miners and the result starts looking like another version of proof of work without much production value.
Ambient makes the opposite design choice. Rather than letting miners run arbitrary models, it standardizes the network around a single large open-weight model and builds consensus around verifying outputs from that model. Where Pearl has miners brute-force the same puzzle, Ambient has them compete through auctions.
Users or agents post inference jobs with a deadline and a price, effectively saying: complete this within X minutes and I will pay Y. Miners bid to take the task. The winner runs the query on the network model and posts a bond that can be slashed for missing the deadline or failing quality guarantees. A randomly chosen validator set then checks the result, with priority weighted by historical useful work rather than by staked capital. The system is built as a Solana fork that replaces stake with useful work and aims to run at Solana-like speed.
The auction mechanism is also central to Ambient’s pricing. A standard API provider has to recover the full cost of serving a request from the user’s payment. Ambient miners can be paid twice for the same unit of work: once by the user or agent that posted the job, and once by the protocol’s reward for verified useful work. In theory, that means miners should bid based on net cost after expected token rewards, passing some of the subsidy through as lower prices. Tcheyan says this is how the protocol tries to turn emissions into cheaper, verified inference rather than into generic mining rewards.
That same structure is meant to solve the issue Pearl leaves open. In Pearl, miners can earn rewards by running matrix multiplications whether or not a real customer asked for the output. In Ambient, miners only earn tokens by winning jobs that someone actually posted and paid for. Mining and serving real demand become the same action by design.

Ambient also uses an unusual approach to output verification. Tcheyan explains that as a language model generates text, each step produces logits, raw scores across all possible next tokens. Those scores act like a fingerprint of the exact model doing the work and can be hashed into a short value. To verify a long output, the validator does not need to rerun the whole task. It picks a random point in the text, asks the miner for the fingerprint at that point, then runs the model for a single token at the same location to check whether the fingerprint matches. According to Ambient, this keeps verification overhead near 0.1%, which it says is roughly 10x to 1000x lower than the overhead of zero-knowledge-proof-based approaches tried elsewhere.
Even with that design, the report says useful proof-of-work systems still face two larger hurdles. The first is demand. Today, the slice of the market willing to pay for verifiable, censorship-resistant, neutral inference that minimizes trust in a provider is still small. If real demand is not there, token rewards alone can fill the network with miners and bring it back to work that is “useful” only in form. The second hurdle is token value capture. Neither Pearl nor Ambient has fully closed the loop where real usage creates token demand, token demand funds security rewards, and those rewards support more usage. Tokens are earned and sold, while the product itself, inference or proof, does not yet require broad token use. Tcheyan suggests the likely end state is that these networks will make their native tokens the default payment rail for inference, but whether long-term organic demand can outrun emissions remains unanswered.
USD.AI and the financing layer beneath inference
Venice tokenizes access. Pearl and Ambient tokenize production. Underneath both sits a different market: financing for the GPUs that make inference possible. Tcheyan presents USD.AI as the clearest example in the report of crypto doing something it is naturally good at, precisely because it does not revolve around creating a new token demand story. Instead, it raises capital in a more conventional way, routes stablecoin deposits into loans for operators buying GPUs, and repays depositors from leasing cash flow.
Large operators already finance equipment through bank facilities, asset-backed securitizations and private credit. CoreWeave’s multi-billion-dollar GPU-backed debt is the case he cites. Smaller neocloud providers have a harder time. They may own hardware and hold contract cash flows that could support a loan, but they lack the balance sheet, treasury operations and lender relationships needed to move quickly. USD.AI lends to that part of the market.
The structure uses two tokens. Depositors mint USDai, a synthetic dollar backed by PayPal’s PYUSD, which is itself backed by U.S. Treasuries and cash. USDai is non-yielding and designed to stay liquid and composable. To earn yield, depositors stake it into sUSDai, whose value rises as the position accrues rewards. Returns come from two sources: interest paid by GPU borrowers on active loans and Treasury yield earned on idle reserves between deployments. With the loan book running at about half the size of reserves, the report says staking yield is about 8%, and the protocol is targeting 10% to 15% as more capital gets deployed.
Tcheyan spends considerable time on the hard part of any hardware-backed loan: enforcement after default. USD.AI previously recorded each financed GPU as an ERC-721 NFT and described it as a legal document of title under Article 7 of the Uniform Commercial Code, with the borrower holding the machine under a bailment arrangement and the NFT serving as collateral. The protocol called that framework CALIBER. It later abandoned the model because it created too much commercial friction.
Following a July 15 update, the NFT now represents the loan record rather than transferring legal claims on the collateral itself. It carries servicing terms and routes repayments on-chain, while enforcement continues to depend on ordinary off-chain loan documentation and the usual operating stack any hardware lender would need: site inspections, installation proof, collateral monitoring, lien filings and cooperation from data centers or custodians. Tcheyan notes that this enforcement path, and the protocol’s broader model, has not gone through a full distressed recovery cycle.
The report also highlights a classic asset-liability mismatch. A liquid token sits on top of loans that amortize over three years. Many RWA credit protocols hide that by promising instant redemptions and then run into trouble under stress; Tcheyan points to USD0++ depegging as one example. USD.AI does not promise immediate exit. Redemptions clear in 30-day windows based on amortized principal and are handled first come, first served. The protocol will not liquidate a performing loan to meet withdrawals. On top of that, it uses a pricing queue inspired by Flashbots MEV-Boost, letting users who want to skip the line bid for priority, with those fees routed to holders who stay in the queue.
Loan terms resemble CMBS structures: 70% to 80% loan-to-value ratios, borrower reserves covering roughly three months of debt service, liquidation after two missed payments, and hardware that is insured, monitored and recoverable through specialized partners.

Tcheyan includes USD.AI in the inference capital market because it connects credit to pricing. Lenders financing GPUs need some way to mark collateral: how quickly hardware depreciates, what it could fetch in a forced sale, what advance rate is safe, how residual value should be hedged. Compute indexes and the futures curves beginning to form provide that reference point. In turn, real credit exposure gives those prices a use beyond speculation.
USD.AI says about 95% of its loan book is backed by long-term offtake contracts rather than spot rentals, which means borrower repayment capacity depends more on counterparties that precommitted to capacity than on day-to-day rental prices. Two more protections sit between that risk and depositors. First, the protocol says each loan is quickly de-risked through 70% to 80% LTV combined with prefunded three-month debt-service reserves, pulling effective LTV down into the low 60s at origination. Second, roughly one-quarter to one-third of loans are repaid in the first year, allowing exposure to be recovered before hardware value falls too far.
Each new loan now also carries impairment insurance from Barkr. Its AI-driven collateral valuation is backed and reinsured by Munich Re. If a borrower defaults and the collateral sells for less than Barkr’s appraised value, the difference is paid to the protocol. Given the 80% LTV cap, USD.AI describes that as full coverage of outstanding debt.
Insurance changes the risk but does not erase it. The report is explicit on that point. Residual value risk is shifted to a stronger counterparty, but the system becomes more dependent on Barkr’s valuation model and on the continued validity of the reinsurance. Neither the coverage nor USD.AI’s off-chain enforcement has been tested through a true wave of defaults. Compared with a few months ago, though, one bad liquidation now hits the insurer before it reaches depositors.
Galaxy’s conclusion: the market is early, and financing is the clearest fit so far
Tcheyan ends by saying that both on-chain and off-chain inference capital markets are still small relative to the broader growth of AI. For on-chain products to scale, they need to show that the advantages they introduce are durable.
He sees those advantages clearly. Tokenized access models such as Venice convert claims on inference into bearer-like assets that can be held, resold, rented or assigned to agents rather than left trapped inside a revocable account subscription. Useful proof-of-work systems such as Pearl and Ambient use token emissions to push inference below market cost and make outputs verifiable enough that buyers do not have to trust providers not to swap models or cut access. Financing protocols such as USD.AI turn illiquid GPU credit into composable on-chain tools that any stablecoin holder can help fund or exit, often faster than in traditional credit markets.
The headwinds are just as clear. Tcheyan argues that no one has yet connected real demand for compute to real demand for crypto tokens. Production networks mint tokens and sell them, using emissions to subsidize below-market inference. Tokenized access markets still trade more on speculation about the issuer than on usage; DIEM is his main example. Financing is the exception because it has real borrowers with real cash flow and does not depend on minting a token purely to bootstrap interest.
His final point is that the strongest edge for on-chain inference capital markets may not be direct competition with incumbents on cheap, large-scale serving of inference. It may lie in capital formation for markets that traditional finance is too slow, too small or too constrained to serve. That, he suggests, is a pattern crypto keeps rediscovering. It often does not win the product, the exchange, the model or the application itself, but it repeatedly becomes a fast way to build the financial layer around those things through pricing, financing, fragmentation and settlement.
Inference is the latest and largest example in that pattern. The market structure that would treat compute as a financial asset, indexes, futures, credit and tokenized capacity, barely exists today. For Tcheyan, that absence is the opportunity. Financing has found the most concrete demand so far, while the rest of the stack remains a bet that the same logic will extend upward as compute itself becomes more financialized. The report closes by saying inference markets may take years to mature, but the financial layer around them is being built now. It also notes that the section on USD.AI was updated on July 15 to reflect changes in how the protocol handles claims on defaulted loans.

