OpenAI says GPT-Red helped harden GPT-5.6 against prompt injection before release

OpenAI says GPT-Red helped harden GPT-5.6 against prompt injection before release

N
News Editor
2026-07-15 20:50:11
OpenAI has unveiled GPT-Red, an internal automated red-teaming system built to uncover weaknesses in its language models and pressure-test them before deployment. The company said the tool was used ahead of GPT-5.6’s release to improve the model’s resistance to prompt injection attacks, a class of exploits that tries to override or manipulate a model’s instructions. According to OpenAI, GPT-Red was trained with self-play reinforcement learning, allowing it to produce increasingly effective attacks while defender models learned to block them. The company said those attacks were then folded into GPT-5.6’s training pipeline. In OpenAI’s internal evaluations, GPT-Red succeeded in 84% of test scenarios, compared with 13% for human red teamers under the same conditions. OpenAI also described a case study in which the system manipulated an autonomous vending machine agent into cutting prices, ordering discounted stock, and canceling another customer’s order before the issue was disclosed and fixed. The company said GPT-Red will remain an internal tool because it includes deliberately developed offensive capabilities.
OpenAIGPT-RedGPT-5.6prompt injectionAI securityred teamingEthereum Foundation

OpenAI has introduced GPT-Red, an automated AI system built to find security weaknesses in its language models. The name comes from cybersecurity red teaming, the practice of intentionally trying to break a system before real attackers can exploit its flaws.

In a post on Wednesday, OpenAI said GPT-Red helped make GPT-5.6 more resistant to prompt injection attacks before the model was deployed.

"As model capabilities grow, safety and alignment must scale with them," OpenAI wrote on X. "Red-teaming is essential, but today’s approaches are difficult to scale, creating a critical bottleneck. GPT-Red is one way we’re addressing it."

Self-play training produced stronger prompt injection attacks

According to OpenAI, GPT-Red was trained through self-play reinforcement learning. The system generated progressively stronger prompt injection attacks while defender models learned how to resist them. OpenAI said those attacks were folded into GPT-5.6’s training process.

The company reported that GPT-Red succeeded in 84% of internal evaluation scenarios. Human red teamers reached 13% in the same tests.

"GPT-Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models," OpenAI wrote. "Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT-Red to continuously find broader and more complex failures."

Case study involved an autonomous vending machine agent

In one case study, OpenAI said the system manipulated an autonomous vending machine agent into lowering prices, ordering discounted inventory, and canceling another customer’s order before the vulnerabilities were disclosed and addressed.

Part of a longer red-teaming push after ChatGPT

GPT-Red follows several years of cybersecurity work by OpenAI after the public launch of ChatGPT.

In 2023, the company launched its OpenAI Red Teaming Network, bringing in outside cybersecurity researchers and domain experts to test ChatGPT and other models for security flaws before release. GPT-Red extends that effort by automating much of the process, using an AI model to generate prompt injection attacks and other adversarial tests at a scale that would be difficult for human researchers alone.

Part of a broader move to use AI to secure AI

OpenAI’s announcement also points to a wider shift toward using AI systems to secure other AI systems.

Earlier this month, the Ethereum Foundation said it had deployed AI agents to red-team critical network infrastructure, uncovering a vulnerability in software used by Ethereum consensus clients. Researchers said AI agents can search larger codebases than humans, though the challenge has shifted from spotting potential bugs to proving which ones can actually be exploited.

GPT-Red will stay internal

OpenAI said GPT-Red will remain an internal tool because it contains intentionally developed offensive capabilities.

"We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today’s models can be used to make tomorrow’s models more robust, aligned, and trustworthy," the company said.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.