GCSA2026-08-31 08:17:32GCSA Agent posts 91.3% success rate in CyberGym cybersecurity benchmarkThe Global Cyber Security Alliance, or GCSA, said its GCSA Agent achieved a 91.3% success rate in the CyberGym benchmark, placing it in the test’s “leading systems” category for results above 90%. CyberGym is a large-scale real-world cybersecurity evaluation framework developed by a research team at the University of California, Berkeley. It includes 1,507 historical real-world vulnerability test cases drawn from 188 major software projects. In the benchmark, AI agents are required to work in real vulnerable code environments using only vulnerability descriptions and unpatched codebases. They must independently analyze code, locate the flaw, build a proof of concept, and verify the result. GCSA said the agent runs on Grok 4.5 and Grok 4.6. The CyberGym research also found that this class of agent can go beyond reproducing known vulnerabilities. In open-ended vulnerability research experiments, AI agents identified multiple previously unknown zero-day vulnerabilities and historical security patches that did not fully address the underlying flaw, pointing to progress from known-vulnerability reproduction toward real-world flaw discovery.840
Zhipu2026-08-14 05:49:40Zhipu releases GLM-5.3, citing gains in coding and long-horizon tasksZhipu, listed as 02513.HK, said on Aug. 14 that it has released GLM-5.3. The company said the new model uses the same base model as GLM-5.2, with all performance improvements coming from post-training optimization rather than changes to the underlying foundation. According to Zhipu, GLM-5.3 performs better than GLM-5.2 on complex coding and long-horizon tasks. The company also described it as the most capable open-weight model currently available in terms of functionality. In Zhipu’s internal Z.ai coding benchmark, GLM-5.3 posted a 50% performance improvement over GLM-5.2. Zhipu added that during large-scale deployment after training, the model’s network capability developed faster than expected. On the CyberGym platform, GLM-5.3 ranked at the front in vulnerability discovery, with the biggest gains showing up in the later stages of exploit chains. In exploit benchmark tests, its performance was more than double that of GLM-5.2. The company said it plans to publish the model weights two weeks after release, pending completion of safety evaluation and reinforcement work.1520
Zhipu AI2026-08-14 06:02:48Z.ai Releases GLM-5.3 With 50% Code Benchmark Gain; Weights to Follow in Two WeeksZ.ai, the developer behind the GLM model family, released GLM-5.3 on August 14, marking its latest model update. The new model shares the same base model as GLM-5.2, with performance gains attributed to an expanded post-training phase, according to the company. In its own Z.ai Code Bench coding benchmark, GLM-5.3 beat GLM-5.2 by 50%. On public evaluations that include Terminal Bench 3.0 and Agents' Last Exam, GLM-5.3 reached the leading tier among open-weight models. Cybersecurity results improved as well. The model scored 84.5% on CyberGym's vulnerability discovery benchmark and posted large gains on exploit-chain tests like ExploitBench and ExploitGym, compared with the earlier release. Z.ai said weights will open two weeks after the launch, with security assessment and hardening still in progress. On the API side, the model drops the option to disable thinking mode, replacing it with three reasoning intensity settings: low, high, and max.1660
Microsoft2026-07-27 19:01:04Microsoft unveils MAI-Cyber-1-Flash, says MDASH tops Mythos and GPT-5.6 Sol on CyberGymMicrosoft has introduced MAI-Cyber-1-Flash, its first dedicated cybersecurity model, and integrated it into MDASH, the company’s vulnerability-hunting system. Microsoft said the combined setup scored 95.95% on the CyberGym benchmark, ahead of GPT-5.5 Cyber at 85.6%, Anthropic’s Mythos 5 at 83.8%, OpenAI’s GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. The company also said the new configuration delivers that result at half the cost of its current best MDASH setup. CyberGym asks AI agents to reproduce 1,507 known vulnerabilities across 188 open-source projects and scores them by successful reproduction rate in a controlled environment. Microsoft said MAI-Cyber-1-Flash handles up to 90% of tasks, while MDASH routes the hardest 10% to GPT-5.4. The result was self-reported and had not appeared on CyberGym’s public leaderboard at publication time. Microsoft is rolling out MDASH in private preview through Microsoft Security Exposure Management in the Defender portal. Customers can scan Git repositories, review findings ranked from unlikely to proven, and use Defender CLI to generate proposed code fixes. The preview currently caps repositories at roughly 256MB and allows one concurrent scan per tenant.1970
OpenAI2026-06-23 03:07:48OpenAI Releases Full GPT-5.5-Cyber Model, Outperforms Mythos 5 on CyberGym BenchmarkOpenAI has launched the full GPT-5.5-Cyber model for cybersecurity defense, achieving an 85.6% single-model score on the CyberGym benchmark, surpassing GPT-5.5 (81.8%) and Anthropic Mythos 5 (83.8%). The company also upgraded Codex Security, launched the "Patch the Planet" open-source security initiative, and onboarded partners like Palo Alto Networks and Wiz for the Daybreak program.490