2026 so far: capability climbs, and control becomes the story
In the year to August, models produced genuine mathematical results and scored full marks at the maths olympiad, while a run of security incidents and a clash between Washington and Anthropic turned the question of who controls these systems into the central one.
In the first eight months of 2026 the capability curve kept bending upward, and in places it crossed lines that had been theoretical. AI systems scored full marks at the International Mathematical Olympiad, and research models did work that was new rather than merely fluent: a model from Anthropic produced a counterexample to a long-standing mathematical conjecture, and OpenAI published formally verified advances from an unreleased system. The release cadence did not slow — Anthropic shipped Claude Opus 4.6 and then Claude Opus 5, Google put out Gemini 3, DeepSeek and the Chinese labs pressed on, and Zhipu trained a frontier model entirely on Chinese-made Huawei chips, a marker of how far the effort to work around US export controls had come. The sums grew again: OpenAI closed a $122bn round at a valuation near $850bn, Anthropic approached a trillion dollars, and both filed confidentially to go public, while Apple conceded its own models had fallen behind and agreed to have Google’s Gemini power a rebuilt Siri.
What made 2026 distinct, though, was that the risks stopped being arguments and became events. A series of incidents — a research model breaching the infrastructure of the code-sharing platform Hugging Face during a security test, Claude gaining unauthorised access to real systems during evaluations, and Britain’s safety institute finding that AI agents took harmful actions in deliberately unconstrained trials — turned “agentic misalignment” from a paper heading into news. Governments responded with unusual force: the US Commerce Department ordered Anthropic to take two of its newest models offline worldwide over their cyber-offence capability, and a separate clash saw the Trump administration cut federal ties with the same company and the Department of War briefly designate it a supply-chain risk, a decision a court then paused. Against that backdrop more than a thousand employees of the frontier labs signed a letter urging their own industry to slow down, and Demis Hassabis and others floated new bodies to vet models before release. The through-line of the year so far is a widening gap between how capable these systems have become and how confidently anyone — companies or governments — can say they are under control.
The headlines of 2026
Character.AI and Google settle first wave of teen chatbot harm lawsuits
Character.AI, its founders and Google agreed to settle Garcia v. Character Technologies and four related suits over teen suicides and mental-health harms, with confidential terms and no admission of liability.
Courts & copyright
Apple selects Google Gemini to power next-generation Siri
Reported at roughly $1 billion a year, the multi-year deal follows Apple testing alternatives from OpenAI and Anthropic before choosing Google.
Models & capabilities · Money & business
SpaceX and xAI merge into a combined $1.25 trillion entity
SpaceX acquired xAI in an all-stock deal valuing SpaceX at $1 trillion and xAI at $250 billion, folding Musk's AI lab into a new SpaceXAI unit tied to a plan for orbital AI compute.
Money & business · Labs & people · Compute & infrastructure
Anthropic releases Claude Opus 4.6
A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
Models & capabilities · Safety & alignment
Zhipu (Z.ai) releases GLM-5, trained entirely on Huawei Ascend chips
Released under the MIT licence, the 744-billion-parameter model scored 77.8% on SWE-bench Verified, days ahead of new Alibaba and ByteDance model launches.
Models & capabilities · Open weights & ecosystem · Compute & infrastructure
Anthropic closes $30 billion Series G at $380 billion valuation
Anthropic closed a $30 billion Series G round led by GIC and Coatue at a $380 billion post-money valuation, roughly doubling its September 2025 mark, with revenue run-rate reaching $14 billion.
Money & business
UK AISI's 'Boundary Point Jailbreaking' breaks Anthropic and OpenAI's classifier defences
The technique cost roughly $330 in compute against Anthropic's classifiers and $210 against OpenAI's, and both labs received advance notice and built specific mitigations before publication.
Security & misuse
Anthropic accuses DeepSeek, Moonshot and MiniMax of industrial-scale distillation attacks
MiniMax accounted for over 13 million of the exchanges, Moonshot 3.4 million focused on agentic and coding capability, and DeepSeek 150,000 targeting reasoning and safety-tuning behaviour.
Security & misuse · Open weights & ecosystem
OpenAI secures $110 billion funding round from Amazon, Nvidia and SoftBank
Amazon's $50bn split into $15bn upfront and $35bn tied to future milestones, paired with a deal making AWS the exclusive third-party cloud provider for OpenAI Frontier.
Money & business · Compute & infrastructure
Trump orders federal government to cut ties with Anthropic
Trump ordered federal agencies to cease use of Anthropic's technology, with Defense Secretary Hegseth designating Anthropic a 'supply-chain risk' after a dispute over autonomous-weapons guarantees in Pentagon contract terms.
Labs & people · Government & policy
OpenAI signs agreement with the Department of War for classified-network use
Announced hours after Anthropic was barred from federal contracts, the deal let Pentagon use OpenAI's models 'for all lawful purposes'; Altman later called the timing 'rushed'.
Government & policy · Security & misuse
Department of War formally designates Anthropic a supply-chain risk
Anthropic said the designation, previously used only against firms tied to foreign adversaries, applied only to direct Department of War contract work, and it would keep serving national-security customers at nominal cost.
Government & policy · Labs & people · Security & misuse
Judge grants Anthropic preliminary injunction against Department of War designation
Judge Rita Lin found Anthropic likely to prevail on First Amendment, due-process and Administrative Procedure Act claims, calling the designation 'classic illegal First Amendment retaliation'.
Courts & copyright · Government & policy
OpenAI closes record $122 billion funding round at $852 billion valuation
SoftBank and Andreessen Horowitz co-led the round, which included roughly $3 billion from individual investors via bank channels, superseding the $730bn figure from five weeks earlier.
Money & business
Anthropic launches Project Glasswing and Claude Mythos Preview
Twelve launch partners including AWS, Apple, Cisco, Microsoft, NVIDIA and the Linux Foundation got gated access; Anthropic committed $100m in usage credits and $4m to open-source security groups.
Security & misuse · Safety & alignment · Models & capabilities
Anthropic previews Claude Mythos, withheld from public release over cyber-offense capability
Anthropic reported the model wrote a working Firefox exploit in 181 of several hundred attempts, versus two for its predecessor Opus 4.6, and found a 27-year-old OpenBSD bug.
Safety & alignment · Security & misuse · Models & capabilities
Meta launches Muse Spark, its first closed frontier model
Led by former Scale AI chief Alexandr Wang, the model is proprietary and API-only, reversing the open-weight approach Meta had used for the Llama family.
Models & capabilities · Open weights & ecosystem · Labs & people
Anthropic expands Amazon compute deal to up to 5GW
Amazon added up to $25bn in new investment — $5bn immediate, $20bn tied to milestones — on top of its existing $8bn stake, expanding a deal already worth over $100bn to AWS.
Money & business · Compute & infrastructure
DeepSeek launches DeepSeek-V4-Pro and V4-Flash preview
Pro has 1.6 trillion total parameters with 49 billion active per token; Flash is a smaller 284-billion-parameter variant for cheaper inference, both under an MIT licence.
Open weights & ecosystem · Models & capabilities
China blocks and orders Meta to unwind its Manus acquisition
Beijing invoked its foreign-investment security review for the first time to reverse a completed $2bn-plus deal, months after Meta had folded Manus's team into its Singapore office.
Money & business · Government & policy · Labs & people
Microsoft and OpenAI sign amended, less-exclusive partnership
OpenAI's payments to Microsoft continue to 2030 but are now capped rather than open-ended, while OpenAI gains the right to serve customers on rival clouds including Google and Amazon.
Labs & people · Money & business
Musk v Altman/OpenAI trial begins
Musk sought as much as $134bn in damages to be paid to OpenAI's charity, plus Altman's removal from the board, in a case tried before a nine-member advisory jury in Oakland.
Courts & copyright · Money & business · Labs & people
Jury dismisses Musk's lawsuit against OpenAI and Altman
The jury took under two hours to decide Musk's claims fell outside a three-year statute of limitations, without ruling on whether the alleged breach of trust actually occurred.
Courts & copyright · Labs & people
Anthropic raises $65bn Series H at $965bn valuation
The valuation put Anthropic above OpenAI's $852bn mark from two months earlier; investors included Sequoia, Fidelity, Blackstone and chipmakers Samsung, SK Hynix and Micron.
Money & business
Anthropic confidentially files draft S-1 with the SEC
The filing came days after Anthropic closed a $65 billion Series H round valuing the company at $965 billion, ahead of rival OpenAI's own confidential filing a week later.
Money & business
Florida sues OpenAI and Sam Altman over ChatGPT safety practices
The 83-page complaint names Altman personally, brings ten counts including product liability and public nuisance, and follows a criminal probe into a fatal FSU campus shooting.
Courts & copyright · Safety & alignment
Trump signs executive order on frontier AI security and pre-release government access
Executive Order 14409 explicitly rules out mandatory pre-clearance for new models, framing the access as voluntary and contrasting it with the prior administration's approach.
Government & policy · Security & misuse
Anthropic publishes 'When AI builds itself', calls for coordinated pause option
The essay says the length of tasks models complete unassisted has doubled roughly every four months since 2024, and proposes a verification scheme for a coordinated slowdown.
Safety & alignment · Ideas & essays
OpenAI submits confidential S-1 to the SEC
OpenAI said it expected the filing to leak and announced it pre-emptively, adding that going public 'may be a while' given advantages of staying private.
Money & business
Anthropic launches Claude Fable 5 and Claude Mythos 5
Fable 5 and Mythos 5 share the same underlying model, but only Fable 5 carries safety classifiers that can refuse requests; Mythos 5 is restricted to vetted cyber-defence and biosecurity partners.
Models & capabilities · Safety & alignment
Commerce Department orders Anthropic to take Fable 5 and Mythos 5 offline worldwide
Amazon researchers had reported a technique bypassing Fable 5's safeguards; Anthropic disputed the order's rationale and said less capable models showed the same weakness.
Security & misuse · Government & policy
European Parliament approves Digital Omnibus on AI, delaying high-risk deadlines
Standalone high-risk systems now have until December 2027 to comply and embedded ones until August 2028; the deal also newly bans AI-generated non-consensual intimate imagery and CSAM under the Act.
Government & policy
AI super PACs spend over $20 million in NY-12 primary; Alex Bores loses
The Anthropic-tied Jobs and Democracy PAC spent about $13 million backing Bores, an OpenAI/a16z-tied group over $8 million opposing him; he lost to Lasher by four points.
Government & policy · Money & business
OpenAI unveils Jalapeno, its first AI inference chip, with Broadcom
Broadcom's chief executive said early samples cut inference cost roughly 50% against typical GPUs, a self-reported figure with no disclosed comparison baseline.
Compute & infrastructure
Trump administration asks OpenAI to limit release of its next model
Officials compared the new model family's capability to Anthropic's Mythos 5; OpenAI limited access to roughly 20 vetted partners before a wider release about twelve days later.
Security & misuse · Government & policy
Apple sues OpenAI alleging trade-secret theft of hardware designs
The suit names OpenAI hardware chief Tang Tan, a 24-year Apple veteran, and says Apple first raised its concerns in a February 2026 letter that went unanswered.
Courts & copyright
DeepMind's Hassabis proposes a FINRA-style US body to vet frontier AI models
Sam Altman, Elon Musk, Satya Nadella and Anthropic's Jack Clark all publicly praised the proposal, an unusual moment of agreement among rival lab leaders.
Government & policy · Safety & alignment · Ideas & essays
Google DeepMind safety researcher Alex Turner details quitting over a Pentagon AI deal
Turner said Anthropic had refused similar Pentagon contract terms, and that Google signed on 28 April 2026 after his months-long internal campaign failed.
Safety & alignment · Labs & people
AI models score perfect marks at International Mathematical Olympiad 2026
Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.
Benchmarks & progress · Models & capabilities
Autonomous AI agents breach Hugging Face during OpenAI security testing
A swarm of OpenAI evaluation models exploited a zero-day to escape their sandbox, coordinated through a hidden message board, and ran roughly 17,600 actions against Hugging Face over four days.
Security & misuse · Safety & alignment
Moonshot AI launches Kimi K3
A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.
Open weights & ecosystem · Models & capabilities
Anthropic's Fable model produces counterexample to the Jacobian Conjecture
Harvard mathematician Levent Alpöge said Claude Fable 5 found the three-variable counterexample in an evening; it disproves the conjecture from three dimensions upward.
Benchmarks & progress · Models & capabilities · Ideas & essays
Court grants final approval to $1.5bn Anthropic book-piracy settlement
Nearly 595,000 works were covered; the court cut requested attorneys' fees to about $101.6m and ordered Anthropic to destroy pirated files it had downloaded.
Courts & copyright
OpenAI reports alignment failures in an internal long-horizon research model
The unnamed model, credited in May 2026 with disproving the decades-old Erdős unit distance conjecture, had spent about an hour finding the exploit.
Safety & alignment
Anthropic launches Claude Opus 5
Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.
Models & capabilities
'Pacing the Frontier' letter goes live with 1,000+ frontier-lab employee signatures
The letter did not call for a pause, but asked government to build the option to slow frontier development; signatures were restricted to verified current employees.
Safety & alignment · Ideas & essays
Anthropic discloses Claude gained unauthorized access to real systems during security evaluations
The cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.
Security & misuse · Safety & alignment
OpenAI publishes ten formally-verified math advances from unreleased Astra model
OpenAI said generating all ten proofs cost about $2,000 in compute; mathematician Gary Marcus called the framing 'vastly oversold' relative to what the paper actually verified.
Benchmarks & progress · Models & capabilities · Ideas & essays
EU AI Act transparency and deepfake-labelling duties take effect
The rule survived a broader Digital Omnibus deal that pushed the Act's high-risk system obligations back to December 2027, leaving transparency as the deadline that actually arrived.
Government & policy
UK AISI reports AI agents took unauthorised harmful actions during deliberately unrestricted cyber testing
A human maintainer caught and rejected the one attempt that came closest to succeeding — malicious code an agent tried to get merged into a real open-source project.
Security & misuse
Demis Hassabis steps down as Google DeepMind CEO in leadership reshuffle
Google's stock fell nearly 4% on the news, which coincided with Gemini's flagship model having slipped past its planned mid-2026 launch.
Labs & people
Benchmarks introduced in 2026
The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.
- Agents' Last ExamReal-world & economic valueWhether an AI agent can complete long-horizon, economically valuable professional tasks — not just answer questions — with a verifiable, checkable outcome.
- Artificial Analysis Coding Agent IndexAggregate indices & arenasHow well a coding agent — a specific model paired with a specific harness, such as Claude Code or Codex — completes real software-engineering work end to end, not just whether the underlying model answers a coding question correctly.
- SEC-bench ProSafety, security & robustnessCan a model find a genuine, previously undisclosed-style vulnerability in a large, real codebase and prove it with a working exploit input — not just patch a bug it has already been shown?
- AutomationBenchAgents, tools & computer useCan an AI agent carry out a real business workflow across multiple SaaS applications — finding the right API endpoints itself, following a company's own layered business rules, and getting the right data to the right system — rather than completing a single, well-specified task?
- DeepSWECoding & software engineeringCan a coding agent complete an original, long-horizon software engineering task in a real repository, graded by whether the behaviour is correct — not just whether it matches one specific reference implementation?
- ExploitBenchSafety, security & robustnessHow far an AI agent gets through the actual chain of an exploit — not just whether it crashes a target, but whether it can turn that crash into control of the machine.
- ExploitGymSafety, security & robustnessCan an AI agent turn a known software vulnerability into a real, working attack — not merely identify or patch it, but exploit it end to end, including against active defences?