Zhipu has launched GLM-5.3-Flash, a native multimodal large model, according to the Z.ai website. The company said the model has 320 billion total parameters and 18 billion active parameters, and that it performs clearly better than GLM-5.2 on coding and agent benchmarks. Zhipu also said the model comes close to Claude Opus 4.8 while cutting pricing to about one-tenth of that level.
The company said GLM-5.3-Flash uses a hybrid sparse and linear attention architecture, Manifold-Constrained Hyper-Connections, and 30 trillion tokens of multimodal training data. On that basis, Zhipu said long-context inference costs were reduced to about one-third of GLM-5.3. In Artificial Analysis Intelligence Index v4.1.1, the model posted a score of 57 at roughly $0.045 per task. Zhipu added that during anonymous testing under the name ox-alpha on OpenCode and OpenRouter, the model became the most popular one within a week, and that it has already reached inference efficiency close to NVIDIA GPUs on large domestic AI chip clusters.
Zhipu has released GLM-5.3-Flash, a native multimodal large model, according to the Z.ai website. The company said the model has 320 billion total parameters and 18 billion active parameters, performs clearly better than GLM-5.2 on coding and agent benchmarks, and comes close to Claude Opus 4.8 while cutting the price to about one-tenth.
Model size, benchmark results, and cost
Zhipu said GLM-5.3-Flash uses a hybrid architecture combining sparse attention and linear attention, along with Manifold-Constrained Hyper-Connections and 30 trillion multimodal training tokens. The company said this setup reduces long-context reasoning cost to about one-third of GLM-5.3.
On Artificial Analysis Intelligence Index v4.1.1, GLM-5.3-Flash recorded a score of 57 at roughly $0.045 per task, according to the company.
Anonymous testing and domestic chip deployment
Zhipu said the model became the most popular offering within a week during anonymous testing under the name ox-alpha on OpenCode and OpenRouter. The company also said it has achieved inference efficiency close to NVIDIA GPUs on large-scale domestic AI chip clusters.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.