OpenAI has released a paper titled Towards safety cases for frontier AI training, saying frontier reinforcement learning training should complete a structured safety document before it proceeds. The company said the ideal form would resemble the "safety case" model used in aviation and nuclear power.
OpenAI added that AI is difficult to govern with the same level of rigor as traditional high-risk industries because capability jumps can produce emergent complexity. It said the framework is still being put into practice internally and is being opened for external feedback.
Scope limited to frontier RL training
OpenAI said the proposal covers only frontier RL training. Internal deployment and external deployment, it noted, still need to be assessed against a broader set of alignment properties.
Three-layer technical structure
On the technical side, OpenAI divides the safety case into three layers: alignment training, containment, and monitoring.
Alignment training
For alignment, the paper calls for both automated and human review of RL environments, tuning reward models to prevent reward hacking, running offline alignment evaluations and backtesting, and preventing automated evaluators from reading chain-of-thought so models do not learn how to evade oversight.
Containment
For containment, OpenAI emphasizes layered sandboxing and infrastructure hardening, red-teaming with training checkpoints, limiting high-bandwidth communication across samples, and preserving conversation records with write-once storage.
Monitoring
On monitoring, OpenAI says models should remain observable and maintain high recall on historical incidents. New risks should trigger additional evaluations, and high-priority alerts should either receive a human response within an agreed time window or automatically pause training.
Operating controls and governance
OpenAI also sets out operating recommendations. Separate teams should write objection pre-mortems. Senior leaders, including the head of research, the head of safety, and the chief scientist, should be able to veto the start of training. The training lead should be accountable for both the safety case and incident response.
If a safety case fails, training should be paused according to a manual. The materials should be disclosed to an internal oversight committee, and auditors should receive enough access to verify the process. Severity tiers should be offset by design, and on-call staff should be allowed to escalate directly to the CEO.
The system should default to fail closed and should not be able to start without monitoring. It should also be able to roll back downstream data and scoring that have been contaminated by misaligned models.
Incident investigations modeled on aviation practice
For investigations, OpenAI points to aviation-style procedures. During an inquiry, the company proposes internal rolling updates, using ablation and resampling to trace training dynamics, and conducting operational and cultural reviews.
It also calls for building detection evaluations that do not directly fit incident samples, then using incident-derived evaluations as regression tests. OpenAI said investigation findings, post-mortems, and process changes should be disclosed publicly after the investigation concludes, and affected third parties should be informed as quickly as possible.

