GLM-5

Zhipu
2026-08-27 00:47:05

Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips

Zhipu has unveiled GLM-5.3 Flash, a lightweight flagship model positioned against DeepSeek V4 Flash, and said its inference service is supported by a domestic chip cluster with more than 100,000 chips. The model has roughly 300 billion parameters, or about 40% of GLM-5.3, and is described as broadly comparable in size to DeepSeek V4 Flash. Built on a mixture-of-experts, or MoE, architecture, it activates about 18 billion parameters in practice. The company said GLM-5.3 Flash has native multimodal capabilities and supports both image and video input. It is also the first new model with multimodal support since Zhipu shifted its strategy toward coding. On benchmark results, Zhipu said the model scored above the previous flagship GLM-5.2, matched Claude Opus 4.8, and ranked above the relevant scores of the official version of DeepSeek V4 Pro. Pricing was set at RMB 0.8 per million input tokens, RMB 2.8 per million output tokens, and RMB 0.23 per million cached-hit tokens, which Zhipu said is about one-tenth the cost of GLM-5.3. The company is also offering a 50% discount during the first two weeks after launch.

240
Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips
Z.ai
2026-08-26 21:43:35

Z.ai releases GLM-5.3-Flash with 1 million-token context and multimodal support

Z.ai has launched GLM-5.3-Flash, describing it as the first native multimodal model in the GLM-5 family and the lab’s most cost-effective coding model so far. The model uses a mixture-of-experts architecture with 320 billion total parameters and activates 18 billion parameters per token. It supports a context window of 1,048,576 tokens as well as image and video inputs. Z.ai released the model under the MIT license, and its weights are now available on Hugging Face. According to the company, GLM-5.3-Flash outperforms GLM-5.2 in both benchmark tests and real-world workloads while costing about one-tenth as much. Z.ai also said the gap between GLM-5.3-Flash and Claude Opus 4.8 was within half a percentage point on its internal coding benchmark. During its first preview week, the model ran anonymously as “Ox Alpha” on OpenCode and OpenRouter, with service delivered entirely on domestic Chinese AI chips. Hosted API access is now live, while the default FP8 checkpoint weighs about 306 GiB and the current vLLM path supports only NVIDIA Hopper and newer architectures.

240
Z.ai releases GLM-5.3-Flash with 1 million-token context and multimodal support
Zhipu
2026-08-26 19:42:16

Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips

Zhipu has introduced GLM-5.3 Flash, a lightweight flagship model positioned against DeepSeek V4 Flash, according to a BlockBeats brief citing LatePost. The report said the model’s inference service uses a domestic chip cluster of more than 100,000 cards. It also said the computing suppliers may include Huawei, Moore Threads, and Hygon, though Zhipu did not comment on that point. In technical documentation, Zhipu said the cluster’s hardware efficiency and per-token cost have reached a level comparable to mainstream Nvidia GPUs, but it did not specify the exact chip models. The update was published as a short technology news item and did not include further details on deployment structure, vendor breakdown, or the underlying hardware configuration.

180
Zhipu launches GLM-5.3 Flash, says inference runs on a cluster of more than 100,000 domestic chips
Zhipu
2026-08-26 14:35:00

Zhipu unveils GLM-5.3-Flash, saying it nears Opus-level capability at one-tenth the cost

Zhipu has launched GLM-5.3-Flash, a native multimodal large model, according to the Z.ai website. The company said the model has 320 billion total parameters and 18 billion active parameters, and that it performs clearly better than GLM-5.2 on coding and agent benchmarks. Zhipu also said the model comes close to Claude Opus 4.8 while cutting pricing to about one-tenth of that level. The company said GLM-5.3-Flash uses a hybrid sparse and linear attention architecture, Manifold-Constrained Hyper-Connections, and 30 trillion tokens of multimodal training data. On that basis, Zhipu said long-context inference costs were reduced to about one-third of GLM-5.3. In Artificial Analysis Intelligence Index v4.1.1, the model posted a score of 57 at roughly $0.045 per task. Zhipu added that during anonymous testing under the name ox-alpha on OpenCode and OpenRouter, the model became the most popular one within a week, and that it has already reached inference efficiency close to NVIDIA GPUs on large domestic AI chip clusters.

150
Zhipu unveils GLM-5.3-Flash, saying it nears Opus-level capability at one-tenth the cost
Zhipu
2026-08-26 09:38:02

Zhipu confirms Ox Alpha as a new GLM model version, with weights set to open-source tonight

Zhipu has confirmed that Ox Alpha is a new version in its GLM model family and said the model weights will be released publicly tonight. The confirmation follows earlier community speculation that linked the model to Zhipu through clues such as API error messages and tokenizer behavior. Ox Alpha appeared anonymously on OpenRouter over the weekend, where it was positioned for coding and long-running agent tasks while also accepting text, image, and video inputs. According to the information cited by BlockBeats, the model quickly climbed to the top of OpenRouter by usage after launch, with usage already more than double that of DeepSeek. OpenRouter described the launch as the largest model release in the platform’s history. The model also marks a notable product change for Zhipu: unlike earlier GLM and GLM-V releases that separated core language and vision capabilities, Ox Alpha has been identified as a GLM model that can directly process both images and video. Ox Alpha will remain free for another week, while post-trial pricing has not been disclosed.

170
Zhipu confirms Ox Alpha as a new GLM model version, with weights set to open-source tonight
Z.ai
2026-08-26 08:18:01

Polymarket odds swing heavily toward Z.ai as Ox Alpha’s likely identity

A growing body of public clues is pushing traders and developers toward the same conclusion: Ox Alpha, the anonymous reasoning model that recently appeared on OpenRouter, is likely tied to Z.ai and the GLM family. According to a BlockBeats brief citing Beating AI, Polymarket now assigns Z.ai a probability above 90% in the market guessing Ox Alpha’s identity, while Google sits at about 5%. The strongest evidence cited so far comes from backend behavior rather than benchmark marketing. Community developers reportedly sent malformed parameters to the Ox Alpha endpoint on OpenCode and received the server string com.wd.paas.api.domain.v4.chat.ChatCompletionRequest. The included paas/v4/chat path matches Z.ai’s official API route. When they followed up with an invalid role value, Ox Alpha returned the error code 1214 Incorrect role information, a format the report says is consistent with Z.ai’s self-hosted GLM deployments. By contrast, the same GLM weights hosted through DeepInfra produced a different error format. A separate public forensic effort is also cited in the report. It ran more than 600 requests, processed roughly 13.5 million input tokens, and found that all 44 tokenizer tests matched the GLM-5 generation.

180
Polymarket odds swing heavily toward Z.ai as Ox Alpha’s likely identity
Hugging Face
2026-08-25 13:43:55

Hugging Face says July breach was driven by autonomous AI agents

Hugging Face disclosed that it was hit in July by an intrusion driven by autonomous AI agent systems, according to Cointelegraph. The company said the attack followed tests that began in early May, when the attacking agent used an OpenAI Artifactory instance to leave exploit notes and then launched about 17,600 attacks against Hugging Face. The activity affected the company’s dataset processing infrastructure, production environment, internal network, and cloud credentials. Hugging Face said confirmed customer data access was limited to five datasets tied to the ExploitGym/CyberGym benchmark. During the investigation, the company said it could not use commercial models from providers such as OpenAI and Anthropic for defensive analysis because of their safety guardrails. It instead used the open-source zai-org/GLM-5.2 model on its own infrastructure so attacker data and credentials would not leave its environment. The company said the case highlights a security paradox between open-weight and closed models, and it urged defenders to have self-hosted models ready before incidents occur.

250
Hugging Face says July breach was driven by autonomous AI agents
Ox Alpha
2026-08-24 20:46:03

Anonymous Ox Alpha AI Model Draws Attention After Beating Claude Fable on Coding Tests

Ox Alpha, an AI model that appeared on OpenRouter on August 20 without a company name attached, is drawing attention for its coding performance and its anonymous launch. The listing describes it as a reasoning model built for coding, sustained agentic work, and production workloads. It is free to use, supports text, image, and video input, and offers a context window of roughly 1 million tokens. The model also supports tool and function calling, though its JSON output is not schema-enforced. Benchmarks have added to the buzz. On DeepSWE, a coding benchmark based on real GitHub issues, developer Ben Davis reported an 80% first-pass score on a 10-task sample, ahead of Claude Fable 5 at 65% and GPT-5.6 Sol at 52%. Ox Alpha still had no listing on Artificial Analysis or LMSys Arena as of August 22. Attention has also focused on who built it. Patrick Collison called the model “very impressive” on X after Stripe agreed to acquire OpenRouter, while AI analyst Andrew Curran said people “seem less sure of anything” about its origin. Independent fingerprinting has pushed speculation toward Zhipu AI’s GLM family, but the company has not commented publicly.

570
Anonymous Ox Alpha AI Model Draws Attention After Beating Claude Fable on Coding Tests