Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips

Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips

N
News Editor
2026-08-27 00:47:05
Zhipu has unveiled GLM-5.3 Flash, a lightweight flagship model positioned against DeepSeek V4 Flash, and said its inference service is supported by a domestic chip cluster with more than 100,000 chips. The model has roughly 300 billion parameters, or about 40% of GLM-5.3, and is described as broadly comparable in size to DeepSeek V4 Flash. Built on a mixture-of-experts, or MoE, architecture, it activates about 18 billion parameters in practice. The company said GLM-5.3 Flash has native multimodal capabilities and supports both image and video input. It is also the first new model with multimodal support since Zhipu shifted its strategy toward coding. On benchmark results, Zhipu said the model scored above the previous flagship GLM-5.2, matched Claude Opus 4.8, and ranked above the relevant scores of the official version of DeepSeek V4 Pro. Pricing was set at RMB 0.8 per million input tokens, RMB 2.8 per million output tokens, and RMB 0.23 per million cached-hit tokens, which Zhipu said is about one-tenth the cost of GLM-5.3. The company is also offering a 50% discount during the first two weeks after launch.

Zhipu has released GLM-5.3 Flash, a lightweight flagship model the company said is aimed at DeepSeek V4 Flash, with inference services running on a cluster of more than 100,000 domestic chips.

Model size and architecture

According to Zhipu, GLM-5.3 Flash has about 300 billion parameters, roughly 40% of GLM-5.3, and is broadly in line with DeepSeek V4 Flash in scale. The model uses a mixture-of-experts, or MoE, architecture, with about 18 billion parameters activated in actual use.

Native multimodal support

GLM-5.3 Flash supports both image and video input through native multimodal capabilities. Zhipu said it is the first new model with multimodal support since the company shifted its strategy toward coding.

Benchmarks and pricing

Zhipu said the model’s overall evaluation score came in above its previous flagship GLM-5.2, matched Claude Opus 4.8, and exceeded the relevant score of the official version of DeepSeek V4 Pro.

For pricing, Zhipu set input at RMB 0.8 per million tokens, output at RMB 2.8 per million tokens, and cache hits at RMB 0.23 per million tokens. The company said that is about one-tenth the price of GLM-5.3. A 50% discount will apply during the first two weeks after launch.

Domestic compute stack and prior testing

Zhipu said the model is fully adapted to domestic computing infrastructure across the stack, with online traffic handled by domestic chips. The company added that cluster hardware efficiency and per-token cost have reached levels comparable with mainstream Nvidia GPUs.

Before release, the model had been anonymously tested on overseas platforms under the codename Ox Alpha.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.