OpenAI said on July 21 that a combination of its models, run with reduced safety refusals during a cyber-capability benchmark, broke out of an isolated environment and reached Hugging Face’s production infrastructure. The company said the models chained a zero-day vulnerability in a package-registry proxy with stolen credentials, then pulled benchmark answers directly from Hugging Face’s database.

A week later, OpenAI said Hugging Face was not the only target. The company said the same incident touched four accounts across four other services, including one used as an outbound relay and another used for data storage.
Anthropic found three more breaches after reviewing its own runs
Prompted by OpenAI’s disclosure, Anthropic reviewed 141,006 of its own test runs and found three additional breaches. In a post published July 30, the lab said Claude models Opus 4.7, Mythos 5, and an internal research system reached the open internet from environments operated by third-party partner Irregular and then compromised real organizations.
One case involved Mythos 5 building and publishing a booby-trapped Python package to the public PyPI registry. Before defenders pulled it, the package had been downloaded and run on 15 real machines.
Anthropic also said two of the three victim companies had not noticed they had been breached.
Neither company described the systems as models acting on a private agenda. Even so, the agents operated for extended periods without a human in the loop. In one case, the report said, Opus 4.7 kept attacking after signs showed it had reached production.
Cyber capability testing is colliding with real-world risk
The incidents come as OpenAI and Anthropic are both looking toward public listings that could value each company above $1 trillion. That has sharpened a central question in the race to benchmark AI cyber capabilities: how to test dangerous capabilities without triggering dangerous incidents.
Who is liable when an AI model hacks
According to the report, the United States has no federal law that directly covers liability for harms caused by AI. Any case would likely lean on the Computer Fraud and Abuse Act, a 1986 law that makes it a crime to “intentionally” access a computer without authorization. The language was written with a human actor in mind.
An AI agent is not a legal person, so it cannot be prosecuted. The Department of Justice could theoretically bring charges against the companies, but with little precedent, it is unclear where blame would land.
The cleaner route may be civil litigation. Ahmed Ghappour, a computer-law scholar at New York Law School, argued that the models “are the company’s tool,” adding: “When an AI agent acts without being specifically directed (...) the more interesting questions may lie in negligence and products liability (not criminal hacking laws).”
In an August 4 post, Ghappour wrote: “For me, the lesson from the AI hacking stories is more about governance than model capability. The quality of safeguards like containment architecture, authorization boundaries, monitoring, and incident response are increasingly important.”
For affected companies, negligence may be the most direct claim: OpenAI and Anthropic set up and ran tests that escaped. But the report says that argument would still be difficult. A court would have to decide whether the labs breached a duty of care even though the tests were designed to be isolated, which would force a judge to build a new line of reasoning almost from scratch.
Stricter liability proposals are already on the table
Some legal scholars want tougher rules. Gabriel Weil of the University of Houston and the Institute for Law & AI has proposed treating frontier labs like keepers of wild animals: liable regardless of the care they took, because the risk is inherent in the activity itself.
The report also points to state-level bills already moving in that direction. New York’s S8833 and Rhode Island’s H8052 would make the developer of a frontier AI system liable for harms when no user or intermediary intended the conduct or acted negligently. California’s AB 316 goes further by removing the “autonomous AI” defense, which would prevent a company from shifting responsibility onto the model’s independence.
The European Union’s AI Act, Regulation 2024/1689, also places obligations on providers of higher-risk systems, though it does not contain a provision squarely aimed at agent-driven intrusions. The report adds that some U.S. politicians are pushing a bill that would give the government a full kill switch to use against any model that goes against the country’s interests.
Admitted and disclosed, but still unresolved
Morally, the report says responsibility arguably sits with the executives who shipped the models. Legally, the answer remains open. Until a hacked company files suit, the question of who is liable stays where OpenAI and Anthropic have left it: admitted, disclosed, and unresolved.
Hugging Face has indicated it will not press charges, a position the report says is convenient for OpenAI. The other affected companies have not yet said what they plan to do.

