OpenAI says Astra may have reached a critical cyber capability threshold, pauses some internal tests

OpenAI says Astra may have reached a critical cyber capability threshold, pauses some internal tests

N
News Editor
2026-08-07 17:05:37
OpenAI said its upcoming Astra model may have crossed, or be close enough that it cannot rule out crossing, the "Critical" cyber capability threshold under the company’s Preparedness Framework. In a safety statement published on Aug. 7, OpenAI said Astra has shown major progress in agentic coding and cybersecurity tasks, prompting the company to halt some internal activities that do not yet meet heightened security requirements. The response includes tighter controls around high-capability models, such as isolated testing environments, restricted access to networks and tools, stronger encryption for model weights, and sandboxed execution. The company also said it has deployed a general monitoring system across Astra’s agentic applications in both training and evaluation, including assessments of chain-of-thought behavior. If dangerous conduct or deviations are detected during high-risk activity, the system is designed to trigger a safety response and stop the operation. OpenAI added that the move is unrelated to the recent Hugging Face exploit incident. The company said it will work with relevant government agencies and selected AI safety organizations to test Astra’s upper-bound capabilities and share recommended controls with third-party testing partners.

OpenAI said its next-generation model, Astra, may have reached the "Critical" threshold for cyber capability under the company’s Preparedness Framework, a development that has led it to pause some internal testing activity and tighten security controls around the model.

The company disclosed the update in an official safety statement published on Aug. 7, titled “Responding to the Next Frontier of Critical Cyber Capabilities.” In that statement, OpenAI said Astra has made significant gains in agentic coding and cybersecurity. After combining those results with expert assessments, the company said it can no longer rule out the possibility that Astra has reached the “Critical” risk threshold, and decided to disclose that assessment publicly in the interest of transparency with the public and the security community.

How OpenAI defines a critical cyber threshold

Under the Preparedness Framework OpenAI published in December 2023, a model reaches the “Critical” threshold only if it shows an extremely high level of autonomous offensive capability. OpenAI said that would mean a model could, without human intervention, identify and develop functional zero-day exploits of varying severity against real-world critical systems. It could also qualify if, after receiving only high-level goal instructions, it could independently design and execute a novel end-to-end cyberattack strategy against heavily defended targets.

For comparison, OpenAI said earlier models such as GPT-5.6-Sol were rated only at the “High” risk level in cyber capability evaluations and did not meet the critical threshold. The company also stressed that Astra remains under benchmark testing and evaluation, so the assessment process is still ongoing.

Some internal activities have been halted

In response, OpenAI said it has suspended internal activities involving Astra that do not yet satisfy strengthened security requirements. It has also imposed tighter controls for high-capability models, including isolated testing environments, restricted access to networks and tools, stronger encryption protections for model weights, and sandboxed execution.

OpenAI said it has also introduced a general monitoring system across all of Astra’s agentic applications, covering both training and evaluation. That system evaluates the model’s chain of thought. If dangerous behavior or deviations are detected during high-risk activity, the system is designed to trigger a safety response immediately and interrupt the operation.

Company says the move is unrelated to the Hugging Face incident

OpenAI separately said the latest safety escalation is unrelated to the recent Hugging Face exploit event. According to the company, the added controls stem from Astra’s own internal evaluation results rather than from any single outside security incident.

OpenAI plans to work with governments and safety groups

Beyond internal safeguards, OpenAI said it will work closely with relevant government agencies and selected AI safety organizations to test Astra’s upper-bound capabilities. It also said it will provide recommended security controls to third-party testing partners.

The company noted that this kind of proactive disclosure has precedent. In June 2025, when one of its models approached a high-capability threshold in biology, OpenAI said it took similar steps by publicly describing the situation and strengthening safeguards.

OpenAI said advanced cyber-capable AI systems should ultimately be used to help defenders identify and fix vulnerabilities before attackers can exploit them. The company said it wants Astra and future frontier models to be deployed under a responsible framework supported by continued cross-sector cooperation and strict security controls.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
10200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.