OpenAI said its next-generation model, Astra, may have reached the "Critical" threshold for cyber capability under the company’s Preparedness Framework, a development that has led it to pause some internal testing activity and tighten security controls around the model.
The company disclosed the update in an official safety statement published on Aug. 7, titled “Responding to the Next Frontier of Critical Cyber Capabilities.” In that statement, OpenAI said Astra has made significant gains in agentic coding and cybersecurity. After combining those results with expert assessments, the company said it can no longer rule out the possibility that Astra has reached the “Critical” risk threshold, and decided to disclose that assessment publicly in the interest of transparency with the public and the security community.
How OpenAI defines a critical cyber threshold
Under the Preparedness Framework OpenAI published in December 2023, a model reaches the “Critical” threshold only if it shows an extremely high level of autonomous offensive capability. OpenAI said that would mean a model could, without human intervention, identify and develop functional zero-day exploits of varying severity against real-world critical systems. It could also qualify if, after receiving only high-level goal instructions, it could independently design and execute a novel end-to-end cyberattack strategy against heavily defended targets.
For comparison, OpenAI said earlier models such as GPT-5.6-Sol were rated only at the “High” risk level in cyber capability evaluations and did not meet the critical threshold. The company also stressed that Astra remains under benchmark testing and evaluation, so the assessment process is still ongoing.
Some internal activities have been halted
In response, OpenAI said it has suspended internal activities involving Astra that do not yet satisfy strengthened security requirements. It has also imposed tighter controls for high-capability models, including isolated testing environments, restricted access to networks and tools, stronger encryption protections for model weights, and sandboxed execution.
OpenAI said it has also introduced a general monitoring system across all of Astra’s agentic applications, covering both training and evaluation. That system evaluates the model’s chain of thought. If dangerous behavior or deviations are detected during high-risk activity, the system is designed to trigger a safety response immediately and interrupt the operation.
Company says the move is unrelated to the Hugging Face incident
OpenAI separately said the latest safety escalation is unrelated to the recent Hugging Face exploit event. According to the company, the added controls stem from Astra’s own internal evaluation results rather than from any single outside security incident.
OpenAI plans to work with governments and safety groups
Beyond internal safeguards, OpenAI said it will work closely with relevant government agencies and selected AI safety organizations to test Astra’s upper-bound capabilities. It also said it will provide recommended security controls to third-party testing partners.
The company noted that this kind of proactive disclosure has precedent. In June 2025, when one of its models approached a high-capability threshold in biology, OpenAI said it took similar steps by publicly describing the situation and strengthening safeguards.
OpenAI said advanced cyber-capable AI systems should ultimately be used to help defenders identify and fix vulnerabilities before attackers can exploit them. The company said it wants Astra and future frontier models to be deployed under a responsible framework supported by continued cross-sector cooperation and strict security controls.

