The Global Cyber Security Alliance, or GCSA, said its GCSA Agent achieved a 91.3% success rate in the CyberGym benchmark, placing it in the test’s “leading systems” category for results above 90%. CyberGym is a large-scale real-world cybersecurity evaluation framework developed by a research team at the University of California, Berkeley. It includes 1,507 historical real-world vulnerability test cases drawn from 188 major software projects. In the benchmark, AI agents are required to work in real vulnerable code environments using only vulnerability descriptions and unpatched codebases. They must independently analyze code, locate the flaw, build a proof of concept, and verify the result. GCSA said the agent runs on Grok 4.5 and Grok 4.6. The CyberGym research also found that this class of agent can go beyond reproducing known vulnerabilities. In open-ended vulnerability research experiments, AI agents identified multiple previously unknown zero-day vulnerabilities and historical security patches that did not fully address the underlying flaw, pointing to progress from known-vulnerability reproduction toward real-world flaw discovery.
GCSA Agent recorded a 91.3% success rate in the CyberGym benchmark, according to an announcement released today by the Global Cyber Security Alliance (GCSA). The result placed the system in CyberGym’s “leading systems” category, which covers entries with success rates above 90%.
CyberGym is a large-scale real-world cybersecurity evaluation framework developed by a research team at the University of California, Berkeley. The framework includes 1,507 historical real-world vulnerability test cases taken from 188 major software projects.
Benchmark requires full vulnerability analysis workflow
Under the benchmark, AI agents operate in real vulnerable code environments and receive only a vulnerability description and an unpatched codebase. They are expected to independently analyze the code, locate the vulnerability, build a proof of concept (PoC), and carry out verification.
GCSA Agent runs on Grok 4.5 and Grok 4.6, according to the announcement. GCSA said the result shows AI capability across an end-to-end security analysis workflow.
CyberGym research points to broader vulnerability research potential
CyberGym’s research also found that the capabilities of this type of agent are not limited to reproducing known vulnerabilities. In open-ended vulnerability research experiments, AI agents identified several previously unknown zero-day vulnerabilities, as well as historical security patches that failed to fully resolve the underlying issue.
The findings, as cited in the report, suggest autonomous vulnerability analysis is moving from reproducing known flaws toward discovering real ones. The item cited BeInCrypto as the source.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.