Kerman Kohli argues AI compute demand remains open-ended despite a 30%-40% semiconductor selloff

Kerman Kohli argues AI compute demand remains open-ended despite a 30%-40% semiconductor selloff

N
News Editor
2026-07-30 04:42:48
Semiconductor and AI-linked stocks have dropped 30% to 40% from their highs, but Kerman Kohli argues the pullback does not settle the debate over whether the buildout has gone too far. In a Substack post cited by BlockTempo, Kohli says the real question is whether demand for compute is finite or effectively unlimited. He points to demand coming from four groups — governments, scientists, enterprises and individuals — and says each has a different willingness to pay for access to compute. A central data point in his case is the rise in hyperscaler compute order backlog from $500 billion to $2 trillion in roughly 1.5 years. He also argues public cloud pricing shows why hyperscalers have strong incentives to keep adding capacity, describing cloud customers as deeply locked into ecosystems such as GCP and AWS. To illustrate pricing, he compares a self-built $15,000 AI inference machine with Google Cloud pricing, estimating a 4.6-month payback period at on-demand rates and about 10 months under a three-year commitment. Kohli also says open-source models do not eliminate the need for high-end hardware, noting that SOTA systems still require substantial memory and serving capacity. He frames AI compute as a recursive category of demand, where useful applications create the need for still more compute, rather than behaving like a conventional infrastructure cycle.

Semiconductor and AI-related stocks have fallen 30% to 40% from their peaks, reviving calls that the trade has already topped out. Kerman Kohli takes the opposite side. In a recent Substack post cited by BlockTempo, he argues the selloff reflects a misread of what matters most: whether demand for compute is capped, or whether it keeps expanding.

His headline point is simple. If hyperscaler compute order backlog has climbed from $500 billion to $2 trillion in about 1.5 years, he says, calling the entire buildout a bubble requires far stronger proof than a sharp drawdown in listed stocks.

The debate, in his view, starts with one question

Kohli reduces the semiconductor and AI investment case to a single test: do you believe demand for compute is finite or infinite? He argues that too many market participants extrapolate from their local experience with enterprise adoption and treat that as a proxy for the whole market.

He does not dismiss the idea that enterprise adoption can take time. But he says that is only one slice of the picture. At a higher level, the current AI buildout is about bringing compute online, and that compute can be consumed by several distinct buyer groups:

  • governments, for military and defense use
  • scientists, for medicine and other frontier research
  • enterprises, to build products and expand without the same labor constraints
  • individuals, for coding, design, creation and question-answering

Those groups do not share the same budget ceiling. Some users may step back if prices rise, but Kohli argues that does not mean no higher-paying buyer will take their place. For sovereign governments and hyperscalers, compute spending can be non-negotiable. For other buyers, it can still be cheaper than scaling with human labor once legal, management and operational costs are factored in.

Why he rejects the “overbuild” argument

Kohli says the idea that the industry is overbuilding rests on a much stronger assumption than it first appears. To believe AI infrastructure has already overshot demand, he argues, you would also have to believe that governments no longer need smarter weapons and defenses, scientists are satisfied with current research output, companies do not want more product growth, and individuals have stopped wanting to do more with better tools.

If someone truly believes those conditions hold, then the bubble thesis follows. If not, he says, the case for a hard ceiling on compute demand becomes much weaker.

That is why he frames the present moment as a broad competition around compute rather than a standard cyclical spending wave.

$2 trillion of backlog and the hyperscaler incentive to keep building

One of the biggest objections to AI capex is that hyperscalers may be overspending and putting future cash generation at risk. Kohli’s answer is to look at the economics of public cloud itself.

He argues that cloud services are extremely expensive and that large providers know how to extract value from customers. In his telling, they have persuaded a generation of companies that scaling without them is unrealistic, while charging 10x to 20x markups on ordinary non-CPU compute and adding fees for logs, egress and several other services needed for routine workloads.

The lock-in effect is central to his argument. Inside GCP or AWS, bandwidth can be cheap. Moving data and workloads out can get expensive very quickly. Because tenant environments are so detached from bare metal, he says, leaving a hyperscaler can turn into a multi-year effort, if it is possible at all. In that setting, telling a customer that GPU capacity is exhausted is not just inconvenient. It risks pushing that customer toward a rival cloud.

For that reason, Kohli argues that insufficient compute is close to existential for hyperscalers. Keeping enough capacity online is about survival and retention, not just growth.

His pricing example: a $15,000 inference machine versus Google Cloud

To show how much room there is in cloud pricing, Kohli points to a machine he previously built for AI inference at a cost of $15,000 and compares it with a similar Google Cloud GPU instance.

Using the Google Cloud accelerator-optimized pricing page cited in the article, he says the on-demand price is about $3,248 per month, while a three-year reserved commitment comes to about $1,444 per month.

He notes that his own machine had 128GB DDR5 memory, while the Google Cloud configuration listed 180GB of memory, though the exact memory type was not specified.

From there he gives two payback estimates:

  • about 4.6 months using on-demand pricing
  • about 10 months using the three-year commitment price

He says the math on other machines looks similar. For H200 clusters released in late 2024, he writes that payback comes in at under two years.

Kohli also says this is only an illustrative comparison. It does not include land, power, financing or on-site labor, so it is not a complete economic model. Even so, he uses it to argue that hyperscalers have substantial pricing power in rented compute.

Older GPUs, memory economics and his “hardware may appreciate” view

Kohli goes further than a standard cloud margin argument. He writes that GPUs from five years ago are both holding value and seeing rental costs rise. In his framework, hardware has two components: compute, which should get cheaper over time, and memory, which can become more valuable.

That split matters to his broader thesis. New chips may improve compute efficiency, but rising memory value could offset some of the expected decline in hardware economics. If that holds, then returns on hyperscaler capex may end up much stronger than many investors currently assume.

He also references a post by Gavin Baker on X as a useful summary of hyperscaler credit-risk concerns, though the summary cited by BlockTempo does not spell out the contents of that post.

Why he treats backlog as the clearest demand signal

Kohli does not try to model every individual customer case. Instead, he says there is a point where investors need to acknowledge the most direct signal available: people are lining up to pay.

According to the article, compute order backlog grew from $500 billion at the start of 2025 to well above $2 trillion within 1.5 years. Citing EpochAI, Kohli says this kind of customer-driven backlog makes it difficult to argue that all of the spending is fake or economically irrational.

He also pushes back on the idea that the number is mostly a laboratory story. Based on the EpochAI framing cited in the piece, frontier labs account for part of global compute demand, but not all of it.

Open-source pressure does not remove the need for top-end hardware

Kohli breaks the lab question into training and inference. On the training side, he says a new state-of-the-art, or SOTA, model can cost hundreds of millions of dollars, which makes it an investment asset expected to generate lifecycle value through inference, even if the depreciation curve is steep.

Open-source models, in his view, do create pressure. They spread into the market and compete with frontier labs at lower cost. But he does not see that as proof that the premium for top models disappears.

His reasoning is that SOTA systems do not solve every problem for every user, yet they can handle tasks that current model classes cannot. That leaves room for a premium where capability truly matters.

He uses Kimi K3 as an example. The point he makes is not just that the model is open source, but that running a model at that tier still requires uncommon hardware. He says Kimi K3 itself needs roughly 1.5TB to 2TB of memory.

Kohli also notes that Kimi ran out of capacity after opening access to K3. A model can be available in principle, but someone still has to serve it. That still requires capable hardware and enough deployed capacity.

From that angle, he expects labs to become more competitive over time, but not to lose their business models outright. Some workloads, he says, will keep paying for frontier-level capabilities, and the ability to supply that capacity remains important in its own right.

Inference margins, labs and where the hardware demand ends up

On inference, Kohli is more direct. He says inference has already proven profitable and puts service-provider margins at 50% to 70%.

Even if users leave hosted providers and buy their own systems, the underlying need does not disappear. The demand still has to land on somebody’s hardware. Because inference is now the dominant workload, he argues, demand remains well above supply.

He does not claim the major labs will enjoy exceptional margins forever. What he rejects is the idea that they simply fail. The article says there is not enough clear data on lab-level margins, but Kohli believes it is reasonable to infer that their inference optimization is already near industry standard. He also points to the fact that firms with tens of millions or even hundreds of millions of monthly active users are not straightforward frauds or Ponzi schemes, and that their revenue is still growing quickly.

Even if returns on SOTA development are not yet compelling, he argues, that does not mean the market will stop rewarding stronger models. Frontier systems represent the next category of problems AI can solve. In his view, betting against the frontier is not the same as being disciplined on valuation; it can also mean accepting a much harder catch-up path later.

The key distinction from railroads or the internet: recursive demand

Many investors think through AI by analogy, comparing it with the internet buildout, railroads or older infrastructure cycles. Kohli says that misses a defining feature of AI: compute can create more demand for compute.

With railroads or the internet, wider adoption and user saturation matter. AI behaves differently in his framework. If a person or organization finds a useful application with positive ROI, more compute can be deployed and turned into another income-generating or cost-saving engine. That is what he means by recursive demand.

He adds that user experience varies widely. Many people still treat AI as a single-prompt question-answering tool. In his own case, he says agents are becoming indispensable as they grow more capable, and that has pushed his own spending on compute higher over time.

An AI trade he calls ignored, crowded and misunderstood at once

Kohli describes the AI trade as one of the most interesting in the market because it is, in his words, ignored, crowded and severely misunderstood at the same time.

He argues that many participants, including large institutional investors, do not fully understand the technical complexity or how it changes balance sheets and bottom lines. Simplified headlines around models such as Kimi K3, he says, have triggered broad selling in memory stocks even though K3 is one of the more memory-intensive models that can be run.

The article also says many of these stocks, despite already rising, still trade at single-digit forward price-to-earnings multiples. In Kohli’s reading, that implies the market is bracing for a sharp hit to profits and demand in 2028 or 2029.

Robotics, population trends and a much longer buildout

Kohli extends the argument beyond current data center demand. He says many compute forecasts do not account for the rise of robotics over the next five years, and he writes that shorting compute is effectively shorting robots.

He links that view to population decline and the limits of older growth models that relied on more workers producing more output. Without AI and robotics, he argues, the economy loses a path to growth and new value creation.

In the same discussion, he says bringing more compute online is necessary to avoid a slowdown in human progress. He writes that North America is already close to its limit for putting compute online, with Canada and Australia next in line, followed by broader global data center construction and, eventually, more data centers in space once terrestrial limits are reached.

That section is clearly a long-range view, not a near-term datapoint. Even there, though, he returns to a market-structure argument: supply is measured closely through fab timelines and earnings calls, while demand is tracked in real time and is still underestimated.

His conclusion remains the same from start to finish

For Kohli, compute is not an ordinary commodity. It is a substitute for labor, a national-security input and a productive asset that can recursively expand its own demand.

That is why he brings the entire discussion back to the first question. If demand for compute is finite, then the current buildout may indeed prove excessive. If demand is open-ended, then the present wave of capex points to a much larger economic shift than the market is pricing in.

On that issue, his position is clear: a 30% to 40% pullback in semiconductor and AI stocks does not prove the boom is over. In his reading, it shows that the market still has not settled the harder question of how large compute demand can become.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
750

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.