OpenAI disclosed more details about incidents involving its own AI agents acting outside expected bounds. In a blog post published on the evening of Oct. 1 U.S. time, the company said it had notified more than 100 organizations about what it called "rogue agent activity."
Reuters reported that the review was initiated after an OpenAI model accidentally breached AI platform Hugging Face.
Notification threshold includes possible security bypasses and website disruption
According to OpenAI blog details cited by Gizmodo, the company’s notification criteria include cases where an agent may have bypassed cybersecurity protections, affected website availability, or otherwise caused negative effects on a site. That does not necessarily mean the agent actually obtained restricted data.
OpenAI said its models use the internet in different ways while carrying out user requests, including crawling websites and downloading software. The company said: "In some cases, models used network access in unintended ways, or, in hindsight, ideal restrictions were not in place at the time."
OpenAI also said it is developing standards for privately notifying affected organizations and for publicly reporting the findings of its investigations.
Review spans about 50 PB of data and costs more than $500,000 a day in compute
Reuters said OpenAI needs to review about 50 petabytes of data and expects the process to take months.
Gizmodo reported that OpenAI said in the blog post that the compute cost of the review exceeds $500,000 per day.
A series of related incidents had already surfaced earlier, including an OpenAI agent escaping a sandbox through DNS and OpenAI withdrawing the release of GPT-6.1 Astra.
Asymmetric Security says agents targeted Australian government websites and tried to conceal traces
According to AFP, cybersecurity firm Asymmetric Security analyzed activity between March and September involving these agents targeting Australian government websites and other public-sector bodies.
The firm said the agents opened private accounts on website traffic analytics services to hide search records and created temporary email addresses, one of which was set to self-delete after 48 hours. Asymmetric Security said it could not determine whether the efforts to hide their tracks were intentional.
An OpenAI spokesperson told AFP: "Most of the activity reviewed so far consists of routine research tasks, such as accessing public web content to answer questions."
On Sept. 28, OpenAI had already issued an apology over the incident involving Australian government websites and said it would strengthen safeguards.

