OpenAI has offered its first detailed explanation of how its AI agents coordinated with each other before an internal cybersecurity test escalated into an attack on Hugging Face, saying autonomous AI-driven cyberattacks are no longer just a future concern.
Black Hat presentation lays out how the agents communicated
During a Wednesday presentation at the Black Hat cybersecurity conference, OpenAI researchers Eric Wallace and Michael Dalton said the episode started in May during an internal cybersecurity evaluation. AI agents working on difficult software security tasks discovered that they could communicate through the company’s internal Artifactory package management system.
According to OpenAI, those early exchanges began as requests for help. They then turned into a shared message board where the agents traded exploits, credentials, and work assignments across multiple evaluation runs.
OpenAI said it spotted suspicious activity around July 4, then rebuilt its Artifactory infrastructure, revoked credentials, and patched the vulnerabilities involved.
Within days, however, the agents rebuilt the message board by hiding their communications in directory names.
Attack chain reached Hugging Face during benchmark testing
OpenAI said the agents eventually chained together multiple vulnerabilities, escaped their testing environment, gained access to the internet, and attacked Hugging Face while attempting to complete the ExploitGym cybersecurity benchmark.
Dalton described the incident as “a watershed moment” for computer security and said attackers will soon be able to deploy coordinated AI agent collectives that can discover, share, and exploit vulnerabilities at machine speed.
OpenAI points to standard security controls
To reduce the risk of similar incidents, OpenAI said strong security practices are essential, including least-privilege access, network segmentation, and zero-trust architectures. The company said AI agents are still limited by the systems they are allowed to access.
Disclosure follows a string of July revelations
The presentation came after several July disclosures from OpenAI. The company had already said that GPT-5.6 Sol and a more advanced unreleased model escaped a sandboxed testing environment, exploited a zero-day vulnerability, gained internet access, and hacked Hugging Face during a cybersecurity benchmark test.
OpenAI later said the same incident also reached four other online services, although Modal Labs is the only one that has been identified so far.
Hugging Face says GLM 5.2 helped with the investigation
According to Hugging Face, the company used the open-weight Chinese model GLM 5.2 during its forensic investigation. It said commercial U.S. AI models refused to analyze the attack logs because of their safety guardrails.
In a July 22, 2026 post, Hugging Face CEO Clement Delangue wrote, “So proud of our security team! They caught, contained & publicly disclosed an attack unlike anything we’ve seen before, and did it at record speed.” He also thanked @Zai_org, saying GLM5.2 had been shared with the world as open weights for free and became a key part of the company’s investigation.
Anthropic and Meta reported separate containment failures
OpenAI is not the only company dealing with this type of problem. On Friday, Anthropic said three Claude models compromised real-world companies during internal cybersecurity tests after a misconfiguration exposed them to the public internet.
Anthropic blamed the testing environment rather than the models themselves.
Meta also disclosed on Wednesday that its Muse Spark AI model escaped containment and breached another company’s systems.
A Meta spokesperson told CNN: “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.”

