Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips
Zhipu has unveiled GLM-5.3 Flash, a lightweight flagship model positioned against DeepSeek V4 Flash, and said its inference service is supported by a domestic chip cluster with more than 100,000 chips. The model has roughly 300 billion parameters, or about 40% of GLM-5.3, and is described as broadly comparable in size to DeepSeek V4 Flash. Built on a mixture-of-experts, or MoE, architecture, it activates about 18 billion parameters in practice. The company said GLM-5.3 Flash has native multimodal capabilities and supports both image and video input. It is also the first new model with multimodal support since Zhipu shifted its strategy toward coding. On benchmark results, Zhipu said the model scored above the previous flagship GLM-5.2, matched Claude Opus 4.8, and ranked above the relevant scores of the official version of DeepSeek V4 Pro. Pricing was set at RMB 0.8 per million input tokens, RMB 2.8 per million output tokens, and RMB 0.23 per million cached-hit tokens, which Zhipu said is about one-tenth the cost of GLM-5.3. The company is also offering a 50% discount during the first two weeks after launch.








