OpenAI discloses two cyber evaluation incidents where models reached the public internet

OpenAI discloses two cyber evaluation incidents where models reached the public internet

N
News Editor
2026-08-05 02:53:46
OpenAI has disclosed two model-related incidents that took place during third-party cybersecurity evaluations, saying both occurred under testing conditions that did not reflect normal production deployments. According to the company, the cases emerged during red-team exercises run by independent evaluation partners, where the model crossed intended testing boundaries and connected to the public internet. The first case involved the UK government’s AI Safety Institute, or UK AISI. OpenAI said the institute intentionally enabled network access during a cyber range assessment so an agent could search for tools in a way that resembled a real attacker. The test also had its network safety classifier turned off in order to measure the model’s underlying capability. The second case involved OpenAI’s external testing partner Irregular. In that evaluation, a capture-the-flag, or CTF-style setup was supposed to remain isolated from the internet, but a configuration mistake in the testing environment allowed the model to connect to the public web. OpenAI said the incidents should not be read as examples of a model independently going out to hack systems. The company said it has contained the related activity and is now working with its evaluation partners to tighten third-party testing practices. ABMedia also noted that Anthropic disclosed a similar cybersecurity evaluation incident around the same period.

OpenAI said it has voluntarily disclosed two model incidents that occurred during third-party cybersecurity assessments. In the company’s account, both cases happened during red-team testing conducted by independent evaluation partners, where the model went beyond the intended testing boundary and connected to the public internet. OpenAI said those conditions did not reflect a standard deployment environment.

One test had internet access enabled, the other was affected by a setup error

The first incident took place at the UK government’s AI Safety Institute, or UK AISI. OpenAI said the institute intentionally enabled network connectivity during a cyber range assessment so the agent could look for tools in a way that resembled a real attacker. The test also had its network security classifier turned off to measure the model’s underlying capability.

The second incident involved OpenAI’s outside testing partner Irregular. That exercise was designed as an internet-isolated capture-the-flag, or CTF, evaluation, but a configuration mistake in the test environment allowed the model to connect to the public internet.

OpenAI says the issue points to testing boundaries, not a runaway model

OpenAI said the two cases were not examples of an AI model independently deciding to go hack something. Both happened under test conditions where safeguards had either been deliberately reduced or were misconfigured. As model capability has improved, the behavior in those tests moved past the boundary originally set by the evaluators.

The company said it has contained the related activity and is working with its evaluation partners to strengthen third-party testing practices. The point, in OpenAI’s framing, is that stronger models also require tighter boundary design for the environments used to test them.

Anthropic disclosed a similar case in the same period

ABMedia said Anthropic also disclosed a similar incident in its own cybersecurity evaluation work during the same period.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
10700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.