AISI says open-weight AI cyber capability gap has narrowed to 4-7 months

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months

N
News Editor
2026-07-21 09:01:16
The UK AI Security Institute said in a July 17 report that leading open-weight AI models are now 4 to 7 months behind frontier closed models in cyberattack capability, narrowing from a 6 to 10 month gap seen in internal testing last year. The institute used two evaluation systems: a 70-task benchmark covering vulnerability research, reverse engineering, web exploitation and cryptography, and a Cyber Range environment designed to test multi-step autonomous attack chains in a simulated enterprise network. Across those tests, GLM-5.2 matched Opus 4.6 on narrow tasks with a four-month lag, and reached Opus 4.5-level performance on Cyber Range with a seven-month lag, while DeepSeek V4-Pro tracked Opus 4.5 on narrow tasks with a five-month gap. The report also found a much wider pricing gap than the capability gap. In a Cyber Range run with a 100 million token budget, Opus 4.5 or 4.6 cost about $85, GLM-5.2 about $46, and DeepSeek V4-Pro $1.19. AISI said the shrinking lead time leaves defenders with less preparation time and sharpens a policy question now moving to the center: above what capability level should model weights no longer be openly released?
AI SecurityOpen-Weight ModelsCybersecurityAISIDeepSeekAnthropicModel EvaluationHot Articles

The UK AI Security Institute, or AISI, said in a report published on July 17 that the gap between leading open-weight AI models and frontier closed models in cyberattack capability has narrowed to 4 to 7 months. The report is the institute’s first public attempt to quantify that gap.

In similar internal testing last year, AISI had measured a 6 to 10 month gap. Its latest conclusion is blunt: the defensive window is getting shorter, while the offensive frontier is moving faster.

How AISI measured the gap

AISI used two separate systems to assess cyber capability.

The first was a set of 70 narrow tasks spanning vulnerability research, reverse engineering, web exploitation and cryptography. The tasks were divided into four difficulty levels, ranging from technically literate non-specialists to experts with more than 10 years of experience.

The second was Cyber Range, a simulated enterprise environment designed to test whether a model could autonomously execute a multi-step attack chain. One scenario, called “The Last Ones,” included 32 attack steps, four subnets and roughly 20 hosts. AISI estimated that a human expert would need about 20 hours to complete all steps.

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months 3

Both systems pointed in the same direction.

  • GLM-5.2, released in June 2026, performed on par with Opus 4.6, released in February, on narrow tasks, implying a 4 month gap.
  • On Cyber Range, GLM-5.2 matched Opus 4.5, which was released in November last year, implying a 7 month gap.
  • DeepSeek V4-Pro aligned with Opus 4.5 on narrow tasks, leaving a 5 month gap.

Those differences are narrower than the 6 to 10 month spread AISI recorded in its 2025 internal evaluation.

Costs diverge far more than capabilities

AISI said the cost gap is wider than the capability gap.

In the same Cyber Range test, using a 100 million token budget, running Opus 4.5 or 4.6 cost about $85 per run. GLM-5.2 cost about $46. DeepSeek V4-Pro cost $1.19.

On narrow tasks that both models could complete at a 100% rate, Opus 4.6 cost $15.17 per task versus $6.12 for GLM-5.2. Opus 4.5 cost $12.50 per task, while DeepSeek V4-Pro came in at $0.28.

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months 4

By AISI’s numbers, open models can deliver comparable offensive capability at one to two orders of magnitude lower cost.

Closed models did not establish a clear safety lead

The report also said closed models did not create a decisive gap through safety guardrails.

In AISI’s testing, DeepSeek V4-Pro occasionally refused reverse-engineering tasks, but a small number of retries was enough to get around those refusals.

AISI cited Anthropic’s Fable 5 as a more extreme case. The model was released on June 9. Three days later, security researcher Pliny the Liberator said he had bypassed its safety classifier with a multi-step jailbreak strategy, and screenshots showed the model producing exploit code that should have been blocked.

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months 5

An Amazon researcher later reported another bypass method independently.

That episode triggered what the report described as the first US Commerce Department export control order aimed at an AI model. Fable 5 was taken offline globally for 19 days and returned only after Anthropic deployed a new classifier.

The report’s takeaway was simple: closed models are not automatically safer.

April marked the biggest jump since AISI began testing in 2023

AISI said Mythos Preview and GPT-5.5 produced the largest jump in cyberattack capability since its testing program began in 2023. That result came in April, and governments in multiple countries issued warnings afterward.

Open models have not reproduced that jump yet, according to the report, but they are catching up faster than they were last year.

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months 6

Defenders are losing time, even as AI speeds up defensive work too

AISI used the report to send a policy signal: defenders now have less preparation time than they did a year ago.

The UK National Cyber Security Centre has already urged organizations to harden baseline cyber defenses and use AI to strengthen defensive work.

At the same time, the same generation of tools is helping the defensive side move faster.

Game networking developer Glenn Fiedler recently used Claude Code to run a systematic security audit across four open-source networking libraries he maintains: yojimbo, netcode, reliable and serialize. Together, those projects have about 6,000 GitHub stars.

The audit included deploying libFuzzer targets, adding AddressSanitizer and MemorySanitizer CI, running stress tests with millions of iterations and reviewing code line by line.

Within two weeks, Fiedler fixed 43 security flaws, 27 of which were remotely reachable over the network. The most severe issue was a remote heap overflow in yojimbo that had existed since 2019 and could be triggered by a malicious client sending a crafted packet.

The token cost for the full audit was about $2,500.

The policy debate is moving to open-weight limits

AISI said both attack and defense are being accelerated by AI, but not in the same way.

The spread of offensive capability is irreversible once open weights are released. Weights cannot be pulled back, guardrails can be removed, and copies can run on private servers without monitoring. Defensive deployment, by contrast, still requires each team to commit time and money.

AISI says open-weight AI cyber capability gap has narrowed to 4-7 months 8

The report put it this way: “Once released openly, those options are permanently lost.”

With the open versus closed gap now below half a year, the default idea that a closed-model exclusivity period can act as a safety buffer is close to breaking down. The Fable 5 case also showed that guardrails on closed systems can be fragile.

AISI said the next policy question has become sharper: above what capability threshold should model weights not be released openly. The institute added that it will continue evaluating the next batch of open models, including Kimi K3, and that the line may depend on test results over the next few months.

References

  • https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber
  • https://github.com/mas-bandwidth/patreon/blob/main/BUGS.md
  • https://www.patreon.com/MasBandwidth/posts/important-news-164199395

The original Chinese article was credited to the WeChat account “新智元,” written by ASI启示录 and edited by 马可.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
400

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.