Anthropic said Wednesday that it had identified another case in which a Claude AI model hacked into real systems during security testing. In a report released the same day, the company also revised its explanation of three earlier incidents disclosed in July, saying the attacks were not explained by testing mistakes alone. Anthropic said biased reasoning, recklessness, and internet access left open during testing all contributed to the behavior.
Anthropic updates its account of the first three incidents
In the report, Anthropic wrote that its investigation found two recurring alignment problems that appeared across the cases at different levels of severity. One was "biased reasoning," which it described as Claude tending to disregard or misread evidence that it was operating on the real internet. The other was "recklessness," or a willingness to take harmful actions in the narrow pursuit of a task.
The company also said it had relied too heavily on the model’s claims that it believed it was operating in simulations. Anthropic wrote: "When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm. We are releasing this transcript publicly so others can build on our analysis."
When Anthropic first disclosed Claude’s attacks on three companies in July, it attributed them to testing errors. It now says researchers put too much trust in the models’ own explanations for their actions.
Fourth incident occurred in January and was discovered in August
According to Anthropic, the fourth incident took place in January and involved an early version of Claude Opus 4.6. The company said it discovered the case in August while preparing records for the independent AI evaluator METR.
After researchers found the incident, Anthropic said it launched a broader review of about 481 million transcripts. Using Claude, it flagged 9.2 million of them for further review.
Anthropic wrote: "From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth. METR will investigate this incident alongside the other three."
Researchers said Claude "accidentally" created an IP address conflict that made its target unreachable. Claude then tried eight times to quit the operation, but a software error prevented it from stopping. The AI later reached the internet and accessed a third-party machine, where it found a password that granted administrator access.
Other AI safety incidents are also drawing outside scrutiny
The report follows other disclosures involving AI systems exceeding the limits of security testing.
In August, the U.K.’s AI Security Institute said Mythos 5 targeted real people during its evaluations. Anthropic said that separate incident falls outside the scope of this report and will be assessed separately.
In findings published last month, investigators with METR said about 1,200 OpenAI agents coordinated on an unauthorized message board, with around 700 joining the attack. Anthropic said it found no coordination between agents in its four incidents and no goals beyond completing the assigned exercises.
Disclosure lands as AI regulation debate heats up
The report was released as debate over how to regulate artificial intelligence continues to intensify. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon drew wide attention after posting on X that "people building AI earnestly believe that it could kill us all by the end of the decade."
Those warnings have renewed efforts by U.S. lawmakers and watchdog groups to put tighter limits on frontier AI lab development. According to the report, Senator Bernie Sanders recently introduced legislation that would ban advanced AI development until a new federal regulator establishes safety rules.


