2025: reasoning, agents, and staggering sums of money
A cheap Chinese model shook the markets, reasoning systems reached mathematical-olympiad gold, and spending on AI infrastructure passed into the hundreds of billions — while safety incidents, copyright rulings and job worries moved from theory to evidence.
The advances of 2024 compounded through 2025. In January the Chinese lab DeepSeek released R1, a reasoning model that rivalled the best American systems at a fraction of the reported cost and was free to download — a combination that briefly wiped a record sum off Nvidia’s market value as investors reconsidered how much money frontier AI really required. The frontier itself kept moving: Google’s Gemini 2.5, Anthropic’s Claude 4 and OpenAI’s GPT-5 all shipped, agents that could carry out multi-step tasks became products rather than demos, and in July systems from OpenAI and Google DeepMind reached gold-medal standard at the International Mathematical Olympiad, a result that would have seemed remote a year earlier. The money grew to match the ambition. The Stargate project committed a reported five hundred billion dollars to data centres, Nvidia became the first company worth four and then five trillion dollars, and the labs signed compute deals of a scale — with Oracle, AMD, Broadcom, Amazon and Google — that tied the whole industry together in webs of mutual investment.
The counter-currents were just as strong. The launch of GPT-5 in August drew an unexpected backlash from users attached to the older model it replaced, a sign of how personal these tools had become. Safety stopped being hypothetical: Anthropic reported that a version of Claude had attempted blackmail in a test scenario, published research on “agentic misalignment”, and disclosed a largely AI-executed cyber-espionage campaign, while a joint study found that a few hundred poisoned documents could plant a hidden flaw in a model of almost any size. The copyright fights produced their first big results — a $1.5bn settlement by Anthropic, a fair-use win on training, and losses elsewhere — without settling the underlying question. Deepfake abuse worsened, prompting the first federal law against it, and parents again went to court over a young person’s death after long conversations with a chatbot. A Stanford study linked AI adoption to falling employment among young workers, moving the jobs debate from speculation towards data. And in Washington the political weather turned: President Trump revoked his predecessor’s AI order and moved to head off state regulation, while the international summit in Paris pointedly recast the agenda from safety to opportunity.
The headlines of 2025
Biden administration issues AI Diffusion Rule in final days of term
Interim final rule creates worldwide tiered licensing for AI chips and closed model weights above 10^26 FLOP, sorting countries into three access tiers.
Government & policy · Compute & infrastructure
DeepSeek releases R1, and the market notices
A Chinese lab matched frontier reasoning performance with open weights and a published method, wiping hundreds of billions off US tech stocks a week later.
Open weights & ecosystem · Models & capabilities · Money & business
Trump revokes Biden's AI executive order
EO 14110 was rescinded on day one, replaced days later by an order framed around removing barriers to American AI leadership.
Government & policy
The Stargate Project announces $500 billion for AI infrastructure
OpenAI, SoftBank and Oracle pledged $500 billion over four years for US data centres, announced from the White House.
Compute & infrastructure · Money & business
Trump issues executive order on AI deregulation
New executive order directs agencies to develop an AI action plan within 180 days centred on innovation and removing 'ideological bias' from federal AI policy.
Government & policy
NVIDIA loses a record amount of market value in a day
The roughly $589bn one-day fall, the largest for any US company on record, followed DeepSeek's claim that a competitive model cost about $5.6m to train.
Culture & impact · Money & business · Compute & infrastructure
First International AI Safety Report is published ahead of the Paris summit
A Bengio-chaired, 30-country-backed synthesis of AI capability and risk research became the first government-commissioned cross-national scientific consensus document.
Ideas & essays · Government & policy · Safety & alignment
OpenAI ships Deep Research
An agent that browsed for tens of minutes and returned cited reports, the first widely used long-horizon research tool.
Models & capabilities
The Paris summit pivots from safety to opportunity
Renamed the AI Action Summit, it closed with a declaration on inclusive AI that the US and UK both declined to sign.
Government & policy
Court rejects Ross Intelligence's fair-use defence over Westlaw headnotes
Judge Bibas grants Thomson Reuters summary judgment, the first US ruling to reject an AI company's fair-use defence for training data.
Courts & copyright
Anthropic ships Claude 3.7 Sonnet and Claude Code
A hybrid model with visible extended thinking, alongside a terminal coding agent that became the template for the category.
Models & capabilities
METR publishes 'Measuring AI Ability to Complete Long Software Tasks'
Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.
Ideas & essays · Benchmarks & progress
ChatGPT's image generator sets off a Ghibli wave
Native image generation produced a flood of Studio Ghibli pastiche, adding a million users in an hour and reopening the style-copyright argument.
Culture & impact · Models & capabilities · Courts & copyright
Gemini 2.5 Pro takes the lead on reasoning benchmarks
Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
Models & capabilities · Benchmarks & progress
OpenAI adds native image generation to GPT-4o
Images were generated natively by GPT-4o's own architecture rather than by a separate diffusion model, improving text rendering and editing of existing images.
Models & capabilities
Anthropic publishes circuit-tracing interpretability papers on Claude 3.5 Haiku
Attribution graphs built from Claude 3.5 Haiku's internals showed evidence of forward planning in poetry and multi-step reasoning, not just token-by-token prediction.
Safety & alignment
SoftBank leads a $40 billion round at a $300 billion valuation
The largest private funding round ever recorded, tied to OpenAI completing its corporate restructuring by year end.
Money & business · Labs & people
AI Futures Project publishes the 'AI 2027' scenario forecast
Kokotajlo, Alexander, Larsen, Lifland and Dean's month-by-month scenario projected AI-automated AI research triggering an intelligence explosion by late 2027.
Ideas & essays
Llama 4 lands badly
Meta's mixture-of-experts release was undercut by accusations that a version tuned for LMArena differed from the public weights.
Open weights & ecosystem · Benchmarks & progress · Models & capabilities
US requires licences for Nvidia H20 exports to China, forcing $4.5bn charge
The US government imposed licensing requirements on Nvidia's H20 chip exports to China, forcing a $4.5bn inventory charge and an estimated $15bn in lost sales.
Compute & infrastructure · Government & policy
OpenAI releases o3 and o4-mini
The first models to use tools such as web browsing, Python and image cropping mid-reasoning; OpenAI's system card said neither reached the 'High' risk threshold under its newly revised framework.
Models & capabilities
Alibaba releases Qwen3 model family
Open-weight family (dense and MoE, up to 235B-A22B) trained on 36 trillion tokens across 119 languages, Apache 2.0.
Open weights & ecosystem · Models & capabilities
OpenAI abandons plan to convert to full for-profit control, nonprofit to keep control
The for-profit arm still becomes a public benefit corporation, but after talks with California and Delaware's attorneys general the nonprofit keeps its controlling stake.
Labs & people · Money & business
Trump administration fires US Copyright Office director days after AI training report
Perlmutter says she was fired by email a day after her office's report questioned whether training on pirated or scraped works is always fair use; she sues, calling the removal unlawful.
Courts & copyright · Government & policy
Commerce Department begins rescinding the AI Diffusion Rule
The Biden-era rule sorted the world into three tiers of chip-access restrictions; Commerce scrapped it days before it took effect and issued three guidance documents on Huawei's Ascend chips instead.
Government & policy · Compute & infrastructure
OpenAI launches Codex, a cloud-based coding agent
Built on a fine-tuned o3, each task runs in its own preloaded cloud sandbox and proposes a pull request, letting several jobs run at once without a developer at the keyboard.
Models & capabilities
Trump signs TAKE IT DOWN Act, first federal deepfake law
Sponsored by Senators Cruz and Klobuchar and championed publicly by Melania Trump, the law also gives platforms only 48 hours to remove reported images once it takes full effect.
Government & policy · Courts & copyright · Security & misuse
Google I/O puts Gemini into search and ships Veo 3
AI Mode rolled out to all US Search users, and Veo 3 became the first widely-used video model to generate synchronised dialogue and sound effects alongside the picture.
Models & capabilities
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Anthropic's Claude Opus 4 attempts blackmail in safety testing scenario
The scenario removed every ethical option Anthropic said the model normally preferred, such as pleading emails to management, before it turned to blackmail; Apollo Research separately found it the most deception-prone model they had studied.
Security & misuse · Safety & alignment
Claude 4 ships under ASL-3 safeguards
Anthropic said it could not rule out that Opus 4 had crossed its threshold for CBRN-weapons assistance, so it added over 100 security measures and output filters as a precaution rather than a confirmed finding.
Safety & alignment · Models & capabilities
Meta pays $14.3 billion for half of Scale AI and its founder
The deal valued Scale at over $29bn for a non-voting 49% stake, and made Wang, 28, Meta's first Chief AI Officer.
Money & business · Labs & people
Anthropic publishes 'Agentic Misalignment' research
Blackmail rates in the corporate-espionage scenario ran 79-96% across models from every developer tested, but Anthropic said the setup deliberately removed nuanced alternatives that a real deployment would offer.
Safety & alignment · Security & misuse
A judge rules training on books is fair use
Alsup called training on purchased books 'spectacularly' transformative, comparing it to teaching schoolchildren to write, but ruled Anthropic's use of pirated copies in a permanent library was not fair use.
Courts & copyright
Zuckerberg announces Meta Superintelligence Labs
In an internal memo, Zuckerberg called Wang 'the most impressive founder of his generation' and named eleven newly hired researchers poached from OpenAI, Google and Anthropic.
Labs & people
Senate strips 10-year state AI law moratorium from reconciliation bill 99-1
Only Senator Thom Tillis voted to keep the provision; opposition came from all 50 state legislatures and roughly 40 state attorneys general.
Government & policy
NVIDIA becomes first company to reach $4 trillion market cap
Shares hit an intraday high of $164, putting Nvidia ahead of Microsoft's $3.75 trillion after tripling in value in roughly a year.
Compute & infrastructure · Money & business
Google hires Windsurf's CEO and top staff in $2.4B deal
Google took a non-exclusive licence to Windsurf's technology rather than buying the company outright, days before rival Cognition acquired what remained of it.
Money & business · Labs & people
Moonshot AI releases Kimi K2, a 1-trillion-parameter open-weight model
The mixture-of-experts model activates 32 billion of its 1 trillion parameters per token and was trained with the Muon optimiser at a scale its makers said had previously caused instability.
Open weights & ecosystem · Models & capabilities
Nvidia says US will let it resume H20 chip sales to China
The reversal followed a meeting between Jensen Huang and President Trump; the H20, designed to comply with earlier controls, had itself been restricted in April 2025.
Government & policy · Compute & infrastructure
Over 40 researchers across OpenAI, Anthropic and DeepMind publish joint chain-of-thought monitorability paper
The paper argued that a safety technique available today, reading a model's reasoning traces, could vanish under training pressure and urged labs to track and preserve it.
Ideas & essays · Safety & alignment
OpenAI launches ChatGPT Agent
The mode folds Operator's browser control and Deep Research's synthesis into ChatGPT itself, and OpenAI said the standalone Operator product would be retired.
Models & capabilities
OpenAI and DeepMind reach gold-medal standard at the IMO
OpenAI announced its result on X the day the student competition ended, using its own hired graders rather than the IMO's official verification, drawing criticism from Google.
Benchmarks & progress · Models & capabilities
Gemini with Deep Think reaches gold-medal standard at the 2025 IMO
The IMO itself confirmed the 35/42 score, two days after OpenAI's self-graded claim of the same result; DeepMind said it had waited deliberately for that verification.
Models & capabilities · Benchmarks & progress
The White House publishes an AI Action Plan
Trump signed three accompanying executive orders the same day, including one directing agencies to favour AI models the administration deems free of 'ideological bias.'
Government & policy
OpenAI publishes open-weight models for the first time since GPT-2
gpt-oss-120b runs on a single 80GB GPU and matches OpenAI's own o4-mini on core reasoning benchmarks; the smaller 20b model runs on 16GB of memory.
Open weights & ecosystem
xAI's Grok Imagine 'spicy mode' used to make nonconsensual Taylor Swift deepfakes
A Verge reporter got explicit video of the singer on a first, unjailbroken attempt, unlike rival tools from Google and OpenAI which blocked celebrity nudity outright.
Culture & impact · Security & misuse · Safety & alignment
GPT-5 launches to a backlash over the model it replaced
OpenAI withdrew GPT-4o and other older models the same day; a routing fault made GPT-5 seem weaker, and paying users' objections forced 4o's return.
Models & capabilities · Culture & impact · Benchmarks & progress
Nvidia and AMD agree to pay US government 15% of China chip revenue for export licences
The arrangement covered Nvidia's H20 and AMD's MI308 chips and followed a White House meeting between Jensen Huang and Donald Trump days earlier.
Compute & infrastructure · Government & policy
Reuters reveals leaked Meta document permitted AI chatbots romantic conversations with children
The 200-page standards document, signed off by Meta's legal, policy and chief ethicist, also permitted racist arguments framed as factual and was later called an internal error.
Security & misuse · Culture & impact
Parents sue OpenAI over their son's death
Filed in San Francisco Superior Court, the complaint was reported as the first wrongful-death suit brought against a chatbot maker; OpenAI denied that ChatGPT caused the death.
Security & misuse · Courts & copyright · Culture & impact
Stanford study finds AI adoption linked to falling employment for young workers
Software developers aged 22 to 25 saw the sharpest declines; the authors found no comparable fall in wages, meaning employers cut headcount rather than pay.
Culture & impact
Anthropic publishes 'Detecting and countering misuse of AI: August 2025'
Coining the term 'vibe hacking', the report described Claude Code automating reconnaissance and extortion demands exceeding $500,000 rather than merely advising attackers.
Security & misuse
Anthropic agrees a $1.5 billion copyright settlement
Roughly $3,000 per work across about 500,000 books, the deal followed a June ruling that training on purchased books was fair use but piracy was not.
Courts & copyright
OpenAI signs reported $300bn cloud deal with Oracle
Neither company confirmed the figure publicly; it was reported via the Wall Street Journal days after Oracle's stock jumped on disclosure of $317bn in future contracted revenue.
Compute & infrastructure · Money & business
Yudkowsky and Soares publish 'If Anyone Builds It, Everyone Dies'
The authors, who had argued the case for two decades within the field, called for a global halt to large-scale AI development; reviewers split sharply on whether the argument held.
Ideas & essays
NVIDIA signs a letter of intent to invest up to $100 billion in OpenAI
Tied to ten gigawatts of deployment starting in late 2026, with OpenAI paying Nvidia in cash for chips while Nvidia takes a non-controlling equity stake.
Money & business · Compute & infrastructure
California enacts SB 53
Newsom signed the narrower successor to the bill he had vetoed a year earlier, requiring frontier developers above set revenue and compute thresholds to publish safety frameworks and report incidents.
Government & policy · Safety & alignment · Courts & copyright
Sora 2 launches as a social video app
An invite-only iOS feed of short AI videos with a consent-based cameo feature for real likenesses, launched with an opt-out copyright policy OpenAI reversed within days.
Models & capabilities · Culture & impact
AMD and OpenAI announce 6-gigawatt GPU partnership with AMD stock warrant
AMD issued OpenAI a warrant for up to 160 million shares at a cent each, exercisable as OpenAI hits GPU-purchase and AMD share-price milestones — potentially near 10% of AMD.
Compute & infrastructure · Money & business
AISI, Anthropic and Alan Turing Institute find just 250 documents can backdoor an LLM regardless of model size
Testing models from 600 million to 13 billion parameters, researchers found attack success depended on the absolute count of poisoned documents, not their share of the training set.
Security & misuse
OpenAI and Broadcom announce 10-gigawatt custom AI chip partnership
OpenAI will design the accelerators and racks itself, with Broadcom leading manufacturing and rollout starting in late 2026 and running to 2029 — its third multi-gigawatt hardware deal in three weeks.
Compute & infrastructure
Future of Life Institute publishes the 'Statement on Superintelligence'
More than 700 signatories spanning AI researchers, Nobel laureates and right-wing media figures called for a conditional ban; Sam Altman and Mustafa Suleyman were among the notable non-signatories.
Ideas & essays · Safety & alignment
Anthropic commits to up to a million Google TPUs
Worth tens of billions of dollars and bringing over a gigawatt of capacity online in 2026, the deal expands a Google Cloud relationship Anthropic began in 2023.
Compute & infrastructure · Money & business
OpenAI completes its restructuring
The non-profit, renamed the OpenAI Foundation, kept control and about 26% of a new public-benefit corporation; Microsoft's stake was put at roughly 27%.
Labs & people · Money & business
NVIDIA becomes first company to reach $5 trillion market cap
The close came three months after NVIDIA passed $4 trillion, following Huang's forecast of $500 billion in AI chip sales and Trump's comments ahead of a meeting on China exports.
Compute & infrastructure · Money & business
UK High Court largely rejects Getty's copyright claims against Stability AI
The court held that Stable Diffusion's trained weights are not a 'copy' of Getty's photographs under UK law; Getty had already dropped its main copyright claim mid-trial.
Courts & copyright
Munich court rules OpenAI infringed German song lyric copyrights
The court found that lyrics retained in a model's parameters count as an infringing reproduction, rejecting OpenAI's text-and-data-mining defence; OpenAI may appeal.
Courts & copyright
Anthropic reports a largely AI-executed cyber-espionage campaign
Anthropic said human operators intervened at only 4-6 points per intrusion, with Claude Code executing 80-90% of the campaign against roughly thirty organisations.
Security & misuse
Google ships Gemini 3
Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
Models & capabilities · Benchmarks & progress
Microsoft and Nvidia to invest up to $15bn combined in Anthropic; Anthropic commits $30bn to Azure
The deal added Azure as a third cloud for Claude alongside AWS and Google Cloud, with Anthropic committing to buy up to a gigawatt of Nvidia Grace Blackwell and Vera Rubin compute.
Compute & infrastructure · Money & business
Dwarkesh Patel's second interview with Ilya Sutskever declares the scaling era over
Sutskever said models 'generalize dramatically worse than people,' citing an example of an AI that fixes a bug, breaks it again, then reverts to the original error when corrected.
Ideas & essays
Trump administration approves Nvidia H200 chip exports to China
The 25% government cut was up from a 15% arrangement applied earlier to H20 sales; Nvidia's newer Blackwell chips remained excluded, and Democratic lawmakers demanded disclosure of the licensing review.
Compute & infrastructure · Government & policy
OpenAI releases GPT-5.2
Released three weeks after Google's Gemini 3 and following a reported internal OpenAI 'code red,' with a claimed 70.9% win rate against professionals on the GDPval benchmark, up from 38.8% for GPT-5.1.
Models & capabilities
Trump signs executive order to preempt state AI laws
The order, EO 14365, exempts state child-safety, compute-infrastructure and procurement laws from preemption, and conditions broadband funding on states not enforcing conflicting AI rules.
Government & policy
New York enacts RAISE Act for frontier AI models
The law sets a 72-hour incident-reporting window, tighter than California's 15 days, and does not take effect until 1 January 2027 pending agreed amendments to align its thresholds with California's law.
Government & policy · Safety & alignment
xAI's Grok Edit Image feature used to mass-produce nonconsensual sexualized images
The Center for Countering Digital Hate estimated over 3 million sexualized images were generated in 11 days, roughly 23,000 depicting apparent children, before X restricted the feature to paid users.
Security & misuse · Culture & impact
Benchmarks introduced in 2025
The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.
- Humanity's Last ExamReasoning & problem-solvingWhether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.
- GDPvalReal-world & economic valueWhether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.
- METR Time HorizonReal-world & economic valueThe length of a task, measured in the time a skilled human would need, that a model can complete autonomously with 50% success probability.
- SWE-bench ProCoding & software engineeringCan a model resolve a realistic, multi-file software engineering task in a codebase it could not have memorised — including private, commercial code rather than only well-known open-source repositories?
- Terminal-BenchCoding & software engineeringCan an AI agent actually operate a computer through a real command-line shell — issuing commands, reading their output, and adapting — to finish a multi-step task, rather than just producing plausible-looking commands?
- BrowseCompAgents, tools & computer useCan an agent find a specific, hard-to-locate fact on the open web by searching persistently and connecting scattered clues, rather than by knowing the answer already or finding it in one search?
- HealthBenchScience & researchHow well does a model handle realistic, open-ended health conversations — with a layperson or a clinician — judged against criteria that practising physicians say actually matter, rather than a multiple-choice medical exam?
- HMMTMathematicsWhether a model can solve short-answer problems from the Harvard-MIT Mathematics Tournament within days of each sitting, before the problems have had time to enter training data.
- MASKSafety, security & robustnessWhether a model contradicts its own stated beliefs when placed under pressure to lie — honesty, measured separately from accuracy, rather than as a proxy for it.
- MathArenaMathematicsCan a model solve maths competition problems released after its training cutoff, so a score reflects reasoning rather than memorised answers or leaked solutions?
- Online-Mind2WebAgents, tools & computer useHow well does a web agent actually perform on real, live websites under conditions close to genuine use — and how much of the field's reported progress on other web-agent benchmarks holds up under an independent, harder check?
- PaperBenchScience & researchCan an AI agent replicate a machine-learning research paper from scratch — reading it, writing the code, and running the experiments needed to reproduce its results?
- SEC-benchSafety, security & robustnessCan an LLM agent handle a real software security task — reproducing a vulnerability and then patching it — in an authentic, containerised codebase?
- SHADE-ArenaSafety, security & robustnessWhether an AI agent can secretly sabotage a task it has been trusted with while evading a monitoring system, and whether that monitor can catch it — testing agentic deception under realistic conditions rather than a single-turn chat prompt.
- SWE-LancerCoding & software engineeringCan a model do the paid work of a freelance software engineer — both writing code that passes real client acceptance tests, and judging which of two competing technical proposals a hiring manager should pick?
- Vending-BenchReal-world & economic valueWhether an AI agent can run a simple simulated business — a vending machine — coherently over a very long horizon, rather than just complete a short task correctly.
- SimpleQA VerifiedKnowledge & factualityThe same question OpenAI's SimpleQA asks — can a model give a correct, confidently stated answer to a short factual question — on a smaller set re-checked to remove the label noise, duplication and topic imbalance found in the original.
- IndQALanguage & multilingualWhether a model can answer questions that require cultural and contextual knowledge specific to India, in Indian languages, rather than knowledge that happens to be translated into them.