UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier models
On a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.
- Security & misuse
- Open weights & ecosystem
- Notable
The UK AI Security Institute published an analysis of how closely open-weight models were tracking closed frontier models on cybersecurity tasks. It tested Zhipu’s GLM-5.2 and DeepSeek V4-Pro against closed comparators including Anthropic’s Opus 4.6, Opus 4.5, Sonnet 4.5 and OpenAI’s GPT-5.3-Codex, using two methods: a 70-task suite spanning vulnerability research, reverse engineering, web exploitation and cryptography at four difficulty levels, and “cyber range” simulations of full network intrusions, including a 32-step corporate-network attack AISI estimated would take a human expert roughly 20 hours.
On the task suite, GLM-5.2 performed comparably to closed models released around four months earlier, and DeepSeek V4-Pro matched Opus 4.5 from roughly five months earlier. On the cyber-range simulations, GLM-5.2 reached performance comparable to Opus 4.5 within under seven months of that model’s release. AISI said this represented a narrowing from 2025, when open-weight models had lagged the closed frontier by six to ten months.
The gap in running cost was large: AISI estimated roughly 100 million tokens of use cost about $85 for Opus 4.5/4.6, against about $46 for GLM-5.2 and roughly $1.19 for DeepSeek V4-Pro. AISI cautioned it had not attempted specific elicitation techniques or optimisation for the open-weight models, which could understate their true capability, that the cyber-range results carried less statistical weight than the task-suite results because of smaller sample sizes, and that the findings were specific to cybersecurity and should not be generalised to other capability domains.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.