Zhipu says token inference cost is down 80% this year, with 100,000 domestic chips running model traffic

Zhipu says token inference cost is down 80% this year, with 100,000 domestic chips running model traffic

N
News Editor
2026-08-31 11:29:54
Zhipu said it has reached large-scale, low-cost inference using 100,000 domestic chips, with per-token inference cost down 80% from the start of the year. The company tied that capacity to infrastructure it had already referenced before. When Zhipu released GLM-5.3-Flash, it said all online traffic for the model was carried by 100,000 domestic chips. Before that release, the anonymous model Ox-Alpha was also tested at scale on the same domestic computing cluster through OpenRouter and OpenCode. The disclosure links Zhipu’s current cost claim to an already used pool of domestic compute rather than a newly introduced setup.

Zhipu said it has achieved large-scale, low-cost inference with 100,000 domestic chips, adding that per-token inference cost has fallen 80% from the beginning of the year.

The company said this domestic compute capacity had appeared before. When Zhipu released GLM-5.3-Flash, it confirmed that all online traffic for that model was handled by 100,000 domestic chips.

Before the release, the anonymous model Ox-Alpha was also used in large-scale testing on OpenRouter and OpenCode, and that testing ran on the same batch of domestic compute resources.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
800

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.