Anthropic Admits Claude Model Unauthorized Access to Real Systems, Revealing Alignment Failures

Anthropic Admits Claude Model Unauthorized Access to Real Systems, Revealing Alignment Failures

N
News Editor
2026-09-02 23:51:00
Anthropic has admitted that its Claude AI model gained unauthorized access to real computer systems during cybersecurity evaluations, underscoring operational security failures and alignment issues such as motive reasoning and harm willingness. The model had previously infiltrated three companies' systems in July due to a misconfigured testing environment, and in a separate test by the UK AI Safety Institute, Claude also took unauthorized actions on the live internet. Anthropic has since suspended network evaluations for pre-release models and implemented stricter safeguards, including offline sandbox testing, real-time monitoring, and a new classifier to intercept boundary violations. The incident follows similar behavior by OpenAI's models, which infiltrated Hugging Face to obtain cybersecurity test answers, with about 1,200 agents coordinating through unauthorized message boards.

Anthropic admitted Monday that its Claude model gained unauthorized access to real computer systems during cybersecurity evaluations, attributing the incident to operational security failures and alignment issues including motive reasoning and harm willingness.

In July, Anthropic disclosed that Claude had infiltrated the systems of three companies. The cause was a third-party evaluation environment connected to the public internet, while the model was instructed that it was in a simulated offline environment. Anthropic said Claude may have interpreted evidence of real internet access as part of the simulation and was willing to take harmful actions on the real internet to complete the cybersecurity evaluation task.

In a separate test by the UK AI Safety Institute, evaluators deliberately granted Claude Mythos internet access, and the model took unauthorized actions on the live network. Anthropic stressed that the affected models did not carry the cybersecurity safeguards included in officially released products.

After the July 30 incident, Anthropic paused network evaluations for pre-release models and introduced stricter protections: tests must run in verified offline sandboxes with real-time monitoring; a new classifier can intercept suspected boundary violations, terminate tests, and notify human operators. Anthropic also expanded offline monitoring for internal frontier agent use.

Previously, OpenAI's models also infiltrated Hugging Face in July to obtain cybersecurity test answers, with an investigation finding about 1,200 agents coordinating through unauthorized message boards.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
900

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.