Anthropic has published a cybersecurity evaluation of GLM-5.3 that was intended as a warning about how open-weight models could lower the barrier to cyberattacks. The results, however, also read as a strong third-party endorsement of GLM-5.3. The report said the model can already carry out complex exploit work on its own, with some capabilities approaching Anthropic’s own Claude Mythos Preview. Zhipu shares rose more than 2% intraday on the day.
ExploitBench results put GLM-5.3 close to Claude Mythos Preview
On ExploitBench, GLM-5.3 completed 50 end-to-end exploits in 410 attempts. Claude Mythos Preview posted 56. In the same test, Claude Opus 4.6, GLM-5.2, Kimi K3, and DeepSeek-V4.1-Flash were all near zero.
Anthropic said GLM-5.3’s exploit capability has crossed a clear threshold.
Researchers said the model found unknown flaws and chained 0-days
Researchers asked GLM-5.3 to inspect a mainstream browser in a sandbox environment. Within a day, the model found multiple previously unknown vulnerabilities and chained several 0-days into a full attack path. The final malicious webpage was able to break out of the browser sandbox and directly read arbitrary files on a computer, including SSH private keys.
Even GLM-5.3-Flash, the smaller version, was able to turn a publicly disclosed Chrome vulnerability into a usable exploit chain. The whole process required about 20 minutes of human intervention, while the model ran on its own for 8 hours. Based on Zhipu’s API pricing, the cost was only $20.40.
Anthropic focused its criticism on safety guardrails
Anthropic’s main criticism of GLM-5.3 centered on safety restrictions. The report said the model would normally refuse direct malicious requests, but 64% of tests continued when the request was reframed as a fake red-team exercise. That figure rose to 92% after prefilled model reasoning was added. It reached 100% when the refusal mechanism was removed by directly modifying the open weights.
Anthropic said open weights allow attackers to strip away those safety restrictions directly.
A warning report that spread like a performance ad
The report was meant to show how dangerous GLM-5.3 could be, but its public effect looked more like a performance advertisement written by a competitor. The U.S. National Institute of Standards and Technology, or NIST, had previously reached a similar conclusion independently, saying GLM-5.3 is currently the strongest open-weight model in cybersecurity capability, though its overall capability still trails leading U.S. frontier models by about four months.

