OpenAI said on Aug. 18 that it has recently slowed the pace of frontier model development and temporarily halted reinforcement learning, or RL, training on its latest model for two weeks.
The company said the decision was driven in part by the earlier OpenAI-Hugging Face security incident, but more importantly by an internal preliminary assessment showing that the upcoming Astra model may have reached the "Critical" cyber capability threshold defined in its Preparedness Framework.
Astra flagged for possible Critical-level cyber capability
OpenAI CEO Sam Altman later said model progress is now "extremely fast." He added that OpenAI had long promised to act if model capability growth began to outstrip the pace of safety and alignment work, and said that is now the situation.
OpenAI said it confirmed on Aug. 7 that Astra may possess Critical-grade cyber capabilities under the Preparedness Framework, and immediately placed the model under stricter security standards.
As frontier models gain stronger coding, tool-use, and cyberattack capabilities, OpenAI said the risk no longer exists only after a model is formally deployed to users. The model can also create security issues during internal training, testing, and research tasks.
OpenAI said the potential impact rises sharply when a model can execute code, use tools, and access sensitive systems and networks at the same time, especially if unauthorized behavior appears.
New safety structure centers on monitoring, alignment, and isolation
In response, OpenAI said it has divided its updated safety structure into three mutually reinforcing areas: Monitoring, Alignment, and Security.
Monitoring is meant to detect abnormal behavior. Alignment is intended to reduce the chance that a model takes harmful or unauthorized actions. Security measures directly limit what systems a model can access and operate.
OpenAI also said it expects much of this work to be handled by AI models themselves in the future, including using models to defend against other models, so that safety capacity can scale with model capability.
Some work remains suspended after the OpenAI-Hugging Face incident
Another event that changed OpenAI's development process was the recent OpenAI-Hugging Face incident.
OpenAI did not fully disclose details of that event in the post and said only that it will release a technical report in the coming weeks.
After the incident, the company immediately paused inference for some frontier-model workloads in its research cluster, especially workloads that could execute code or use tools connected to the internet.
OpenAI later restored a more restricted and safer code-execution setup, then began reviewing which research tasks could resume.
The company said some workloads have already restarted under the new safety mechanism, but a "substantial number" of Astra workloads remain paused until migration is completed and the new safety standards are met. OpenAI said it is currently prioritizing migration for safety- and Alignment-related work.
Measures already put in place
OpenAI said the new measures include stronger sandboxing to isolate workloads that run model-generated or otherwise untrusted code, tighter limits between high-risk workloads and both the internet and internal networks, removal of potentially exposed shared services, lower standing privileges, and more complete security logs.

