UK AISI publishes first Frontier AI Trends Report
Universal jailbreaks were still found for every system tested, though one model took 40 times more expert effort to break than a predecessor released six months earlier.
- Benchmarks & progress
- Safety & alignment
- Security & misuse
- Notable
The UK AI Security Institute published its first Frontier AI Trends Report, aggregating two years of government evaluations — running since November 2023 across more than 30 frontier systems — using auto-graded tasks, longer-form evaluations, agent simulations, expert red-teaming and human-impact studies spanning cyber, chemical and biological, autonomy and societal domains.
On cyber capability, success on apprentice-level tasks rose from roughly 10% in early 2024 to an average of 50% in 2025, and models completed an expert-level task, requiring judgment associated with ten or more years of professional experience, for the first time that year; the length of task models can reliably complete has also been doubling roughly every eight months. On self-replication, measured with the RepliBench evaluation, success rose from under 5% in early 2023 to over 60% by summer 2025, though models remained comparatively weak at the later, persistence-oriented stages rather than the earlier stage of simply acquiring compute. In chemistry and biology, some models now exceed PhD-level expert performance by up to 60%.
The report also tracked the gap between open-weight and closed frontier models, finding it had narrowed to somewhere between four and eight months depending on methodology — external benchmarks diverge on the precise figure. On security, AISI said universal jailbreaks — prompts that reliably defeat a model’s safety training across many types of harmful request — were found for every single system it tested, though the effort required was uneven: one model took roughly 40 times more expert effort to jailbreak than a predecessor released only six months earlier, while progress in other cases was slower or reversed.
Beyond capability and security metrics, the report cited survey findings on public use of AI: 33% of UK citizens said they had used AI for emotional support in the previous year, and 32% of chatbot users said they had researched election-related topics before the UK’s most recent general election. Taken together, the report functions as AISI’s own retrospective benchmark of the pace and unevenness of frontier progress it has tracked since its founding, rather than a single new finding.