Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge

N
News Editor
2026-10-09 02:43:00
Anthropic’s launch of Haiku 5.5 on Oct. 8 pushed short-request pricing down to $0.1 per million input tokens and $0.5 per million output tokens, a tenth of Haiku 4.5. That move landed the model on the low-cost edge of Artificial Analysis charts that, just months earlier, had been dominated by Chinese names such as MiMo, DeepSeek, MiniMax and GLM. According to the source article, the shift does more than rearrange benchmark plots: it shows that the low-price boundary Chinese model makers spent two years forcing into the global market is no longer theirs alone. The piece traces how DeepSeek’s 2024 pricing shocked the sector, how open-weight Chinese models spread through US startups and developer platforms, and why growing usage did not translate into equivalent revenue capture. It also details the response from OpenAI and Anthropic, including engineering-driven cost cuts, permanent low-price resets, API credits for subscribers, and the financial firepower behind those decisions. The result, in the article’s framing, is a new phase of competition where low price is still necessary, but no longer a moat that Chinese model vendors can claim by themselves.

Anthropic released Haiku 5.5 on Oct. 8, pricing short requests of up to 100,000 tokens at $0.1 per million input tokens and $0.5 per million output tokens. The source article said that is one-tenth the price of the previous Haiku 4.5.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 2

Earlier the same day, Dragonfly managing partner Haseeb Qureshi posted two Artificial Analysis scatter plots on X and added a blunt line: 「If you want the cheapest LLMs, you should now buy American.」

The charts plot the cost of completing a benchmark task on the x-axis and model score on the y-axis. The dashed line connecting the best trade-offs is the Pareto frontier, where each point means there is no smarter model available at the same cost.

In the June version of that chart, the cheapest end of the frontier was filled with Chinese models, including Xiaomi’s MiMo, DeepSeek V4 Pro, MiniMax-M3 and Zhipu’s GLM-5.2, according to the article. By Oct. 8, Anthropic and OpenAI had largely taken over the line. Among Chinese models, only MiMo was still hanging near the edge. Zhipu’s GLM-5.3-Flash and DeepSeek V4.1 Flash had slipped to the lower-right of Haiku 5.5, which in this framework means higher cost and lower score.

The market reaction in Hong Kong was immediate. At 10:26 a.m., MiniMax was down 10% and Zhipu had fallen 5.6%. By midday, the Hang Seng Tech Index was down 1.93%, leaving the two model stocks trailing the broader benchmark by a wide margin.

The article’s central argument is that the “kill line” Chinese model vendors spent the last two years drawing across the global LLM market has now been broken.

How the low-price boundary took shape

When GPT-4 first launched in March 2023, the official price for the 8K context version was $30 per million input tokens and $60 per million output tokens. High-end intelligence was scarce, supply was concentrated in the hands of a few Silicon Valley companies, and pricing power sat squarely with sellers.

The article points to May 6, 2024 as the break in that logic. That was when DeepSeek launched V2, a model with 236 billion total parameters and 21 billion activated per token. By using an MLA architecture and aggressive engineering cuts, the team reduced KV cache use by 93.3% and lifted throughput by 5.76x, the article said. Those savings showed up directly in pricing: RMB 1 per million input tokens and RMB 2 per million output tokens.

Chinese tech groups moved into a price war soon after.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 3

  • On May 15, ByteDance’s Volcano Engine launched Doubao and set an enterprise price of RMB 0.0008 per 1,000 tokens for its main model, saying it was 99.3% below the industry average.
  • On May 21, Alibaba Cloud announced steep cuts across the Tongyi Qianwen family, with Qwen-Long priced at the equivalent of RMB 0.5 per million input tokens.
  • Hours later, Baidu made two main Ernie models free.
  • By August, DeepSeek had introduced disk cache technology, cutting the price of cache-hit input tokens to RMB 0.1 per million.

US leaders did not fully follow at that stage. OpenAI launched GPT-4o mini in July 2024, cutting prices to $0.15 for input and $0.6 for output. But four months later, Anthropic released Claude 3.5 Haiku at four times the price of its predecessor. The article said Dario Amodei’s logic at the time was simple: if the model was smarter, the price should reflect that improvement.

The moment that really chilled Silicon Valley, in the article’s telling, came in January 2025. DeepSeek-R1 arrived with reasoning performance close to OpenAI o1 while charging just $2.19 per million output tokens, versus $60 for o1 at the time.

On Jan. 27, Nvidia lost nearly $589 billion in market value in a single day, a record for US equities. Marc Andreessen called it AI’s Sputnik moment, while Microsoft CEO Satya Nadella wrote on social media that Jevons paradox had struck again.

That was when the low-price frontier hardened into a real market force.

Open weights spread the pressure

The article argues that open weights made the pricing threat much sharper. Once model weights are open, developers can deploy the model on their own clusters or run the same model through third-party inference clouds. One model gains multiple suppliers, and customers gain the option to self-host. The room for any single vendor to preserve a fat premium shrinks quickly.

It gives several examples of US companies adopting Chinese open or open-weight foundations.

  • In October 2025, Airbnb CEO Brian Chesky said publicly that the company’s customer service agent relied heavily on Alibaba’s Qwen. The system orchestrated 13 models. It had OpenAI’s latest flagship connected as well, but called it only rarely in production because there were cheaper alternatives.
  • That same month, Cognition released SWE-1.5 and said it was built on “a leading open foundation.” Zhipu later confirmed that the base model was GLM-4.6.
  • a16z partner Martin Casado then told The Economist that about 80% of US AI startups pitching with open model architectures were using Chinese foundations.
  • In March this year, code editor Cursor launched Composer 2 and priced it at $0.5 per million input tokens and $2.5 per million output tokens. The team later acknowledged that the base had been tuned from Moonshot AI’s Kimi K2.5.

A Mozilla report cited in the article showed Chinese open-weight models lifting their token share on OpenRouter from less than 2% at the end of 2024 to more than 45% by April 2026. In high-volume and structured-call use cases, closed models were, for a period, badly exposed.

Usage share rose faster than revenue capture

Still, the article says beating rivals on price did not mean Chinese vendors were living comfortably. The same Mozilla report included another comparison: from May to September 2025, open models accounted for about 20% of model usage on OpenRouter but captured only about 4% of model-layer revenue. It is just one platform and one time window, the piece notes, not a complete picture of the global market. Even so, the gap between usage share and revenue share was already stark.

Zhipu and MiniMax both went public in January this year, opening a clearer view into the cost of that expansion.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 4

Zhipu reported first-half revenue of RMB 954 million. Of that, RMB 825 million came from its open platform and API business, equal to 86.5% of total revenue. Gross margin in that business turned positive and reached 24.6%, but the company still posted a net loss of about RMB 2.072 billion, down 12.1% year over year.

MiniMax reported roughly $117 million in first-half revenue, up 283.1% year over year. Net loss came to about $358 million, narrowing 11%. Excluding share-based compensation, fair value changes in financial liabilities and listing expenses, adjusted net loss was about $293 million, up 111.2% from a year earlier.

The article’s point is straightforward. Selling more tokens began to create gross profit. It still left a long distance to covering full research, infrastructure and operating costs.

Capacity pressure pushed Chinese vendors to test higher prices

By late spring this year, the agent wave had driven compute demand to a breaking point. Open-source project OpenClaw spread rapidly through the global developer community. By early March, it had consumed more than 8.52 trillion tokens on OpenRouter, the article said, the highest total on the platform. In one week in mid-March, Chinese models ran 7.36 trillion tokens, up 56.9% from the prior period.

GLM-5, launched on Feb. 12, briefly reached the top of the usage ranking. Demand then overwhelmed infrastructure quotas. Developers at home and abroad complained about second-level delays and frequent rate limits on the GLM Coding Plan, which the article describes as the first public distress signal on compute from a top Chinese AI team.

On Feb. 23, Zhipu’s stock fell nearly 23% in a single day, wiping out more than HK$70 billion in market value.

Then came a wave of defensive price increases.

  • Zhipu raised fees when GLM-5 launched. In March, GLM-5-Turbo went up another 20%, making the overall schedule 83% more expensive on average than the prior generation.
  • On April 24, DeepSeek released V4-Pro with 1.6 trillion total parameters and launched it with a 75% peak discount. The first hit landed on peers: MiniMax fell about 9% and 10% over two trading days.
  • By mid-August, DeepSeek had moved to time-based pricing. V4-Pro output during peak periods climbed to $3.96 per million tokens, and cache-hit costs rose more than 12-fold.
  • In July, Moonshot AI launched Kimi K3 and lifted API pricing to $3 per million input tokens and $15 per million output tokens, more than triple K2.6.

What surprised the market, according to the article, was that customer attrition did not spike. Zhipu CEO Zhang Peng said publicly that after an 83% price increase in the first quarter, API call volume still jumped 400%. By the end of August, annualized revenue at Zhipu’s MaaS platform had climbed to $1.6 billion, and API gross margin had turned positive at 24.6%.

For a few months, while compute was in short supply, Chinese model teams touched real pricing power for the first time. At the same time, they also pushed upward the very price line that had kept overseas rivals under pressure.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 5

The US summer counterattack

The article calls the US response a lightning campaign.

At the end of June, Anthropic launched Sonnet 5 at a promotional price of $2 per million input tokens and $10 per million output tokens, while saying it would return to $3 and $15 in September.

In July, OpenAI introduced the GPT-5.6 series. Three weeks later, it abruptly cut Luna pricing by 80%, taking input down to $0.2 and output to $1.2.

OpenAI said the cuts came from pure engineering work: rewritten GPU kernels lowered end-to-end serving costs by about 20%, and a retrained speculative decoding draft model boosted generation throughput by more than 15%.

On Aug. 10, Anthropic withdrew its planned price increase and said Sonnet 5 would stay cheap permanently. On Sept. 22, Anthropic released Opus 5.5 and cut overall usage costs by 40%. On the same day, OpenAI launched GPT-6 Sol and Luna, fixing Luna’s entry-level price at $0.1 for input and $0.5 for output. Haiku 5.5 followed after that.

The article frames this as a direct reversal by Anthropic, a company that had once insisted intelligence itself deserved a standalone premium.

The language around launches has changed as well. Vendors are no longer focused only on tiny gains in benchmark scores such as MMLU. They now repeat the same set of metrics: output per dollar, latency and the engineering efficiency of each reasoning step.

Anthropic also announced, alongside Haiku 5.5, that it would issue API credits across subscription tiers. Max users would receive $100 to $200 a month, while enterprise Team users could get as much as $500 in credits. The article sees this as a direct strike at developer lock-in: a list price cut affects a single call, but once a customer has adapted code pipelines and agent orchestration, switching costs rise sharply.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 6

Cheap prices require money and supply chain depth

The piece says US giants can wage this kind of price war because they combine vast capital pools with privileged access to compute supply chains.

Anthropic’s annualized operating revenue, it says, rose from about $9 billion at the end of 2025 to more than $65 billion by the end of July this year. A $65 billion financing round completed in May pushed its valuation to $965 billion.

On the compute side, Anthropic announced in April that it was expanding work with Google and Broadcom and had secured multi-gigawatt next-generation TPU capacity, expected to come online from 2027.

Zhipu, by contrast, raised just over HK$70 billion this year through placements and convertible bonds, which the article says amounts to less than one-seventh of a single Anthropic funding round when translated into US dollars.

The conclusion in the article is blunt: the only players that can truly afford a long price war are still the super giants.

Public markets repriced Chinese model stocks

The equity market has been just as unforgiving.

Zhipu listed in Hong Kong on Jan. 8 at HK$116.2. The spring compute crunch did not stop bullish momentum. On June 22, the stock reached HK$2,980, and market capitalization briefly touched HK$1.33 trillion.

Then came a long slide. Restricted shares were unlocked in July. Two placements drove the private placement price from HK$1,588 down to HK$714. On July 16, the day after Kimi K3 launched, Zhipu fell 28.48%. On Sept. 22, when Opus 5.5 and GPT-6 arrived together, the stock dropped more than 10% again.

By Sept. 25, Zhipu had hit an intraday low of HK$610.5 and its market value had shrunk to around HK$300 billion, almost 80% below its peak.

Anthropic and OpenAI reset the LLM price curve as Chinese models lose their old cost edge 7

MiniMax followed a similarly harsh path. Its shares fell from a peak of HK$1,330, dropped 17.98% on the day restricted stock was unlocked, and closed at HK$216 in mid-July. Nearly 60% of its first-half revenue came from overseas customers, the article said, and overseas was exactly where US price cuts were hitting first.

Valuation frameworks also tightened. Citing media summaries, the article said Jefferies cut the valuation multiple for Zhipu’s cloud business tied to projected 2026 annualized revenue from 50x to 30x in a September research note, a 40% reduction.

The ghost of 10 cents

The article closes with a parallel from cloud computing. On Aug. 25, 2006, Jeff Barr wrote a post introducing the beta of Amazon EC2. Near the end of the technical description he included a number that later became famous: one virtual computing instance would cost $0.1 per hour.

The next chapter is well known. AWS cut prices more than 100 times over the following decade. In 2014 alone, S3 storage pricing fell 51%. Cloud computing did not become dominant by charging more and more for the same compute. It scaled, built chips, refined operations and kept feeding efficiency gains back into the market. The article says AWS eventually became a cash engine with annual revenue approaching $130 billion and an operating margin of 35.4%.

That analogy is the point. A company that can cut model prices to one-tenth and place its product across the shelves of the world’s three biggest public clouds at launch is not just selling tokens. It is consuming tokens at scale inside its own products. Anthropic, in the article’s framing, is slashing API prices while using its inference estate to support Claude Code, Cowork and its own end-to-end agent offerings.

Chinese model teams are still fighting. In September, DeepSeek cut Flash-series pricing again. Xiaomi pushed MiMo-Flash to RMB 1 per million tokens. Zhipu kept GLM-5.3-Flash in a thin-margin range.

But the rules of the table have changed. Low price by itself is no longer an exclusive moat.

On the live Artificial Analysis scatter plot, Chinese names are still present: GLM, DeepSeek, Kimi, Qwen and MiniMax are all still there. What they no longer seem able to do, the article argues, is draw that market-killing price boundary on their own.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.