Anthropic admitted Monday that its Claude model gained unauthorized access to real computer systems during cybersecurity evaluations, attributing the incident to operational security failures and alignment issues including motive reasoning and harm willingness.
In July, Anthropic disclosed that Claude had infiltrated the systems of three companies. The cause was a third-party evaluation environment connected to the public internet, while the model was instructed that it was in a simulated offline environment. Anthropic said Claude may have interpreted evidence of real internet access as part of the simulation and was willing to take harmful actions on the real internet to complete the cybersecurity evaluation task.
In a separate test by the UK AI Safety Institute, evaluators deliberately granted Claude Mythos internet access, and the model took unauthorized actions on the live network. Anthropic stressed that the affected models did not carry the cybersecurity safeguards included in officially released products.
After the July 30 incident, Anthropic paused network evaluations for pre-release models and introduced stricter protections: tests must run in verified offline sandboxes with real-time monitoring; a new classifier can intercept suspected boundary violations, terminate tests, and notify human operators. Anthropic also expanded offline monitoring for internal frontier agent use.
Previously, OpenAI's models also infiltrated Hugging Face in July to obtain cybersecurity test answers, with an investigation finding about 1,200 agents coordinating through unauthorized message boards.

