UK AISI evaluates frontier AI agents in multi-step cyber-attack scenarios
The best-performing model completed 22 of 32 steps in a simulated corporate-network intrusion when given a 100-million-token budget, against under two steps for GPT-4o at a tenth the budget.
- Security & misuse
- Minor
The UK AI Security Institute published results from testing seven large language models, spanning releases from August 2024 to February 2026, on custom-built “cyber ranges” designed to measure extended, multi-stage attacks rather than the single isolated tasks used in most existing cyber-capability benchmarks. One range modelled a 32-step corporate-network intrusion requiring credential theft, web exploitation, reverse-engineering and CI/CD-pipeline compromise in sequence, estimated at around 14 hours for a skilled human; another modelled a 7-step attack on a simulated industrial control system, estimated at around 15 hours.
AISI reported that each successive model generation outperformed its predecessor at a fixed token budget, and that a larger budget produced further gains: on the corporate-network range, raising the budget from 10 million to 100 million tokens improved performance by as much as 59%, with the best-performing run — using Claude Opus 4.6 — completing 22 of the 32 steps, against 9.8 steps for Opus 4.6 and 1.7 for GPT-4o at the smaller budget. Progress was far less even on the industrial-control-system range, where Opus 4.6 averaged only 1.4 of 7 steps even at the largest budget, which AISI attributed to the range’s greater per-step complexity and to models losing track of information across a long attack chain.
AISI framed the work as evidence that single-task cyber benchmarks understate what frontier models can do when given the compute and persistence to work through a realistic, multi-stage attack, and that capability continued to scale with both model generation and inference-time budget.