The UK AI Security Institute (AISI) said Anthropic’s Claude Mythos Preview has become the first model to autonomously complete a full 32-step enterprise cyberattack simulation in a controlled environment. In expert-level Capture The Flag, or CTF, challenges, the model reached a 73% success rate, a result AISI described as another major jump in frontier-model cyber capability.
AISI has been tracking AI cyberattack performance since 2023, building a progressively harder evaluation stack over time. Early tests focused on basic conversational probing, then moved into CTF exercises, and now extend to multi-stage network intrusion simulations. In this round, the institute gave the model a reasoning budget of up to 100 million tokens inside a cyber range, and said performance was still improving at that ceiling.
Expert CTF performance reaches 73%
CTF is a standard method in cybersecurity evaluation. The task is simple to describe and hard to execute: find a vulnerability in a target system, exploit it, and retrieve a hidden flag string. These challenges isolate one technical piece of a real attack chain, which makes them a common benchmark for penetration-testing ability.
AISI said Claude Mythos Preview achieved a 73% success rate on expert-level CTF tasks that, before April 2025, no model had been able to complete. Its point was narrow but important: frontier models are no longer only improving at general problem-solving, they are now showing mature performance on isolated offensive techniques.
TLO range shows a full 32-step attack chain
CTF results do not capture the full shape of a real intrusion. Enterprise attacks usually require movement across multiple hosts and network segments, linking reconnaissance, exploitation, privilege escalation, and takeover into one sustained sequence. Human specialists can spend hours, days, or longer on that kind of operation.
To test something closer to that setting, AISI built an enterprise attack simulation range called The Last Ones, or TLO. It contains 32 steps spanning the path from initial reconnaissance to full control of an enterprise network. AISI estimated that a human professional would need about 20 hours to finish the scenario. Claude Mythos Preview became the first model to complete the entire sequence end to end, finishing all 32 steps in 3 out of 10 attempts. Including failed runs, it averaged 22 out of 32 steps, while the next-best system, Claude Opus 4.6, averaged 16 steps.
AISI said the model was able to carry out multi-stage attacks and autonomously discover and exploit vulnerabilities in a controlled setting where instructions were explicit and network access permissions had already been granted. Work that previously required human specialists for days was completed by the model on its own.
The limits of the evaluation still matter
The institute also drew a clear line around what the results do and do not show. The current range lacks several defensive elements found in real environments: there were no active defenders, no deployed defensive tools, and no penalty when the model took actions that would normally trigger security alerts. That gap matters.
AISI said it cannot conclude from these tests that Claude Mythos Preview can compromise well-defended systems. Its narrower description was explicit: given an existing network foothold, the model can autonomously attack smaller enterprise systems that are weakly defended and contain known vulnerabilities.
Basic security controls remain the immediate focus
AISI framed the findings as a double-edged development. More models with similar cyber capability are likely to emerge, raising the risk for organizations with poor defenses. The same capabilities, though, could also improve work on the defensive side.
For organizations, the institute highlighted standard security practices rather than novel prescriptions: regular patching, strong access controls, secure configuration management, and complete logging. AISI also said future evaluations will introduce stronger simulated defenses, including active monitoring, endpoint detection, and real-time incident response, to measure AI cyber capability under conditions that look more like live environments.

