Z.ai releases GLM-5.3-Flash with 1 million-token context and multimodal support

Z.ai releases GLM-5.3-Flash with 1 million-token context and multimodal support

N
News Editor
2026-08-26 21:43:35
Z.ai has launched GLM-5.3-Flash, describing it as the first native multimodal model in the GLM-5 family and the lab’s most cost-effective coding model so far. The model uses a mixture-of-experts architecture with 320 billion total parameters and activates 18 billion parameters per token. It supports a context window of 1,048,576 tokens as well as image and video inputs. Z.ai released the model under the MIT license, and its weights are now available on Hugging Face. According to the company, GLM-5.3-Flash outperforms GLM-5.2 in both benchmark tests and real-world workloads while costing about one-tenth as much. Z.ai also said the gap between GLM-5.3-Flash and Claude Opus 4.8 was within half a percentage point on its internal coding benchmark. During its first preview week, the model ran anonymously as “Ox Alpha” on OpenCode and OpenRouter, with service delivered entirely on domestic Chinese AI chips. Hosted API access is now live, while the default FP8 checkpoint weighs about 306 GiB and the current vLLM path supports only NVIDIA Hopper and newer architectures.

Z.ai has released GLM-5.3-Flash, calling it the first native multimodal model in the GLM-5 series and the lab’s most cost-effective coding model to date.

The model uses a mixture-of-experts architecture with 320 billion total parameters and 18 billion activated parameters per token. It comes with a 1,048,576-token context window and accepts both image and video inputs.

Z.ai released GLM-5.3-Flash under the MIT license, and the model weights have been uploaded to Hugging Face. The company said the new model performs better than GLM-5.2 in both benchmark tests and real workloads, while pricing is about one-tenth of GLM-5.2. In Z.ai’s internal coding benchmark, the gap between GLM-5.3-Flash and Claude Opus 4.8 was within half a percentage point.

Z.ai also said the model’s first-week preview ran anonymously under the name “Ox Alpha” on OpenCode and OpenRouter. That preview service was delivered entirely on domestic Chinese AI chips.

The weights are available on Hugging Face, and a hosted API has also been launched. The default FP8 checkpoint weighs about 306 GiB. Z.ai said the current vLLM path supports only NVIDIA Hopper and newer architectures, which means medium-sized and large organizations, as well as AI-native startups renting GPU capacity, can self-host the model.

The update was cited by MarkTechPost.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
2400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.