OpenAI has introduced GPT-Red, an automated AI system built to find security weaknesses in its language models. The name comes from cybersecurity red teaming, the practice of intentionally trying to break a system before real attackers can exploit its flaws.
In a post on Wednesday, OpenAI said GPT-Red helped make GPT-5.6 more resistant to prompt injection attacks before the model was deployed.
"As model capabilities grow, safety and alignment must scale with them," OpenAI wrote on X. "Red-teaming is essential, but today’s approaches are difficult to scale, creating a critical bottleneck. GPT-Red is one way we’re addressing it."
Self-play training produced stronger prompt injection attacks
According to OpenAI, GPT-Red was trained through self-play reinforcement learning. The system generated progressively stronger prompt injection attacks while defender models learned how to resist them. OpenAI said those attacks were folded into GPT-5.6’s training process.
The company reported that GPT-Red succeeded in 84% of internal evaluation scenarios. Human red teamers reached 13% in the same tests.
"GPT-Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models," OpenAI wrote. "Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT-Red to continuously find broader and more complex failures."
Case study involved an autonomous vending machine agent
In one case study, OpenAI said the system manipulated an autonomous vending machine agent into lowering prices, ordering discounted inventory, and canceling another customer’s order before the vulnerabilities were disclosed and addressed.
Part of a longer red-teaming push after ChatGPT
GPT-Red follows several years of cybersecurity work by OpenAI after the public launch of ChatGPT.
In 2023, the company launched its OpenAI Red Teaming Network, bringing in outside cybersecurity researchers and domain experts to test ChatGPT and other models for security flaws before release. GPT-Red extends that effort by automating much of the process, using an AI model to generate prompt injection attacks and other adversarial tests at a scale that would be difficult for human researchers alone.
Part of a broader move to use AI to secure AI
OpenAI’s announcement also points to a wider shift toward using AI systems to secure other AI systems.
Earlier this month, the Ethereum Foundation said it had deployed AI agents to red-team critical network infrastructure, uncovering a vulnerability in software used by Ethereum consensus clients. Researchers said AI agents can search larger codebases than humans, though the challenge has shifted from spotting potential bugs to proving which ones can actually be exploited.
GPT-Red will stay internal
OpenAI said GPT-Red will remain an internal tool because it contains intentionally developed offensive capabilities.
"We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today’s models can be used to make tomorrow’s models more robust, aligned, and trustworthy," the company said.

