2024: reasoning, recognition, and real-world harm
Models learned to think for longer, AI's scientific work won two Nobel Prizes, and the EU's rulebook came into force — as deepfakes reached elections, chatbot-harm lawsuits began, and California's safety bill was vetoed.
Capability grew along new axes in 2024. GPT-4o brought fast, spoken, multimodal conversation; Anthropic’s Claude 3 and 3.5 and Google’s Gemini traded the lead through the year, with Gemini stretching to a million-token context window. The most consequential shift came in September, when OpenAI’s o1 showed that letting a model spend more time “thinking” at the moment of answering — rather than simply making it bigger — could sharply improve its reasoning, opening a second path to progress just as pre-training gains looked harder to come by. AI’s scientific standing was formally recognised: Geoffrey Hinton and John Hopfield shared the Nobel Prize in Physics for the foundations of neural networks, and Demis Hassabis and John Jumper shared the Chemistry prize for AlphaFold, whose third version now modelled interactions across proteins, DNA and drugs. The first genuine agents appeared, too, as Claude gained the ability to use a computer and Anthropic proposed a shared standard for connecting models to external tools.
The harms grew more concrete alongside the capabilities. Deepfakes moved from novelty to weapon: a fabricated robocall imitating President Biden told New Hampshire voters to stay home, and a video call with AI-generated colleagues was used to steal twenty-five million dollars from one company. In one of the year’s most sobering cases, a mother sued the chatbot service Character.AI after her teenage son’s death, beginning a wave of litigation over the effects of these systems on vulnerable users. On governance, the limits of ambition showed: California’s SB 1047, which would have imposed safety obligations on the largest models, was vetoed in September after heavy industry lobbying, and at OpenAI the dissolution of its “superalignment” team and a run of senior departures raised doubts about whether the leading labs would fund safety at the pace they funded scale. The EU AI Act entered into force, the first broad, binding regime — though its real effects lay in the years ahead.
The headlines of 2024
Deepfake video call used to steal $25 million from Arup's Hong Kong office
A finance employee made 15 wire transfers after a video call in which every other participant was AI-generated; Hong Kong police confirmed the case a week later.
Security & misuse · Culture & impact
Anthropic shows backdoored models surviving safety training
Models trained to write secure code unless told the year was 2024 kept the hidden behaviour through supervised fine-tuning, reinforcement learning and adversarial training.
Safety & alignment
AI-generated fake Biden robocall urges New Hampshire voters to skip primary
Thousands of New Hampshire Democrats received a cloned Biden voice telling them a primary vote would forfeit their choice in November; the operative behind it was fined and criminally charged.
Culture & impact · Security & misuse · Government & policy
Microsoft and OpenAI disrupt state-affiliated hacking groups misusing LLMs
The groups used LLMs mainly for reconnaissance, translation, debugging and phishing-content drafting rather than novel attack techniques; the identified accounts were terminated.
Security & misuse
Gemini 1.5 Pro ships a million-token context window
A mixture-of-experts model that matched Gemini 1.0 Ultra on many tasks at lower compute, offered in limited preview with up to a million tokens of context.
Models & capabilities
OpenAI previews Sora, a text-to-video model
Clips up to a minute long held their subjects consistent across camera moves; OpenAI showed samples but gave no public access for ten months.
Models & capabilities · Culture & impact
Musk sues OpenAI over its for-profit turn
The complaint alleged a founding pact to develop AGI 'for the benefit of humanity' had been broken for Microsoft's profit; OpenAI said no such agreement existed.
Courts & copyright · Labs & people
Anthropic's Claude 3 takes the frontier from GPT-4
The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
Models & capabilities · Labs & people
The European Parliament passes the AI Act
Adopted 523 votes to 46 after three years of negotiation; the Council's formal sign-off and publication in the Official Journal followed months later.
Government & policy
NVIDIA announces the Blackwell architecture
The GB200 NVL72 system was claimed to cut LLM inference cost and energy per token by up to 25 times against Hopper-generation hardware.
Compute & infrastructure
Microsoft absorbs Inflection's team without buying the company
Reported terms put the deal at $650 million — most of it a licence fee for Inflection's models, the rest for waiving legal claims over the mass hire.
Labs & people · Money & business
Meta releases Llama 3 and puts its assistant everywhere
8B and 70B open-weight models shipped alongside a much larger, still-training 400B+ version, as Meta AI rolled out across Facebook, Instagram, WhatsApp and Messenger.
Open weights & ecosystem · Models & capabilities
AlphaFold 3 predicts structures across proteins, DNA, RNA and ligands
Restricted at launch to a rate-limited web server rather than downloadable code, prompting an open letter with more than 650 signatures within a week.
Models & capabilities
OpenAI launches GPT-4o with real-time voice
A single model handling text, vision and audio end to end, with conversational latency — and a voice that led to a public dispute with Scarlett Johansson.
Models & capabilities · Culture & impact
Google unveils Project Astra, a universal AI assistant prototype
A prototype, not a product: Google showed a phone-camera assistant with conversational-speed responses but gave no release date beyond 'later this year'.
Models & capabilities
DeepMind publishes the Frontier Safety Framework
A set of internal capability thresholds across autonomy, cybersecurity, biosecurity and ML R&D, joining similar voluntary policies already published by Anthropic and OpenAI.
Safety & alignment
Jan Leike resigns and the superalignment team dissolves
Leike said his team had been 'sailing against the wind' for compute and access; OpenAI reassigned remaining members rather than replacing the team's leadership.
Labs & people · Safety & alignment · Ideas & essays
EU AI Act formally adopted by the Council
The Council's sign-off was the final legislative step; the text still needed formal signature, Official Journal publication, and a staggered two-year rollout before most rules applied.
Government & policy
Anthropic maps millions of concepts inside a production model
Sparse autoencoders extracted human-interpretable features from a deployed model — and turning one up produced Golden Gate Claude.
Safety & alignment · Ideas & essays
Leopold Aschenbrenner publishes 'Situational Awareness'
Aschenbrenner, dismissed from OpenAI's superalignment team months earlier for allegedly leaking information, argued the firing itself illustrated the security failures he described.
Ideas & essays
Apple announces Apple Intelligence with OpenAI inside
ChatGPT access was free and optional, required no account, and Apple said queries would not be logged — with other AI providers to be added later.
Models & capabilities · Money & business
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
Meta releases Llama 3.1 405B
Trained on over 15 trillion tokens with 16,000 H100 GPUs; Meta reported it competitive with GPT-4, GPT-4o and Claude 3.5 Sonnet, though the licence still barred some commercial uses.
Open weights & ecosystem · Models & capabilities
AlphaProof and AlphaGeometry 2 reach silver-medal standard at the IMO
The systems scored 28 of 42 points, one short of gold, but took up to three days on some problems against the competition's 4.5-hour limit.
Models & capabilities · Benchmarks & progress
The EU AI Act enters into force
The world's first comprehensive AI law took effect, banning some uses outright and imposing obligations on general-purpose models by capability.
Government & policy · Courts & copyright
Google takes Character.AI's founders back
The deal, reported at $2.5–2.7 billion, licensed Character.AI's technology to Google without buying the company, and drew Justice Department scrutiny.
Money & business · Labs & people
Scaling test-time compute paper argues extra inference compute can beat bigger models
On some problems, extra inference-time computation matched the gains from a pretrained model roughly 14 times larger, the authors reported.
Ideas & essays · Benchmarks & progress
OpenAI releases o1, trading inference time for reasoning
A model trained to think before answering opened a second scaling axis: spend more compute at inference and accuracy rises.
Models & capabilities · Benchmarks & progress
Constellation Energy to restart Three Mile Island for Microsoft AI power deal
The Pennsylvania reactor, shut since 2019, is due back online in 2028 under a 20-year contract requiring NRC approval, supplying about 835MW.
Compute & infrastructure
Mira Murati and two research leaders quit OpenAI on the same day
Murati cited wanting time for her own exploration; she went on to found Thinking Machines Lab, and Altman called the departures independent and amicable.
Labs & people
Reuters reports OpenAI plans to convert to a for-profit structure
Report says Altman would receive a stake for the first time and the nonprofit board would lose its controlling role, a plan not finalised until late 2025.
Labs & people · Money & business
Newsom vetoes California's SB 1047
The bill would have required safety protocols and shutdown capability for models above a compute threshold; Newsom said it regulated size rather than risk.
Government & policy
OpenAI raises $6.6 billion at a $157 billion valuation
Investors could claw back their money if OpenAI failed to convert to a for-profit structure within two years, and were reportedly asked not to fund Anthropic or xAI.
Money & business · Labs & people
Hinton and Hopfield win the Nobel Prize in Physics
The citation credited Hopfield's 1980s associative-memory network and Hinton's Boltzmann machine, work from decades before the current deep-learning boom.
Culture & impact · Ideas & essays
Hassabis and Jumper share the Nobel Prize in Chemistry
Half the prize went to Baker for computational protein design; the other half was split between Hassabis and Jumper for AlphaFold's structure prediction.
Culture & impact
Anthropic publishes Dario Amodei essay 'Machines of Loving Grace'
The roughly 14,000-word essay argued that a decade of scientific progress could be compressed into five to ten years, while stressing this was an upside scenario, not a forecast.
Ideas & essays
A mother sues Character.AI after her son's death
The complaint sought damages for wrongful death and product liability; Character.AI called the death tragic and said it had since added self-harm safeguards for users.
Courts & copyright · Culture & impact · Security & misuse
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities
OpenAI launches ChatGPT Search
Initially limited to paying subscribers and testers, the feature cited sources inline and drew on licensing deals with Reuters, the Financial Times and other publishers.
Models & capabilities
Anthropic publishes the Model Context Protocol
Anthropic open-sourced the specification and pre-built connectors for tools like Google Drive and GitHub; OpenAI adopted the same standard the following March.
Open weights & ecosystem
BIS issues third major round of chip export controls, adds 140 entities
New rules restricted high-bandwidth memory chips and 24 categories of chipmaking equipment, building on rules issued in October 2022, October 2023 and April 2024.
Government & policy · Compute & infrastructure
Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'
In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
Ideas & essays · Safety & alignment · Security & misuse
OpenAI releases Sora publicly
Every clip carried a visible watermark and embedded C2PA provenance metadata; access launched in the US and Canada only, excluding the UK and EU.
Models & capabilities
o3 posts a breakthrough score on ARC-AGI
A low-compute configuration scored 75.7%, roughly matching the ARC Prize's human-performance threshold, at about $26 per task against roughly $5 for a human solver.
Benchmarks & progress · Models & capabilities
OpenAI announces o3 and opens early access for safety testing
Reported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.
Models & capabilities · Benchmarks & progress
DeepSeek releases V3
DeepSeek's technical report put the final training run at 2.79 million H800 GPU-hours, or about $5.6 million at an assumed $2-per-hour rental rate.
Open weights & ecosystem · Models & capabilities
Benchmarks introduced in 2024
The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.
- FrontierMathMathematicsCan a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?
- AIMEMathematicsWhether a model can solve short, single-answer competition-maths problems requiring several steps of reasoning but no written proof.
- AndroidWorldAgents, tools & computer useCan an agent operate a real Android phone — navigating apps, typing, tapping, adjusting settings — to complete a task described in natural language, with success checked against the device's actual resulting state?
- Artificial Analysis Intelligence IndexAggregate indices & arenasA single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.
- HarmBenchSafety, security & robustnessHow reliably a model refuses to help with a harmful request when an attacker is actively trying to jailbreak it, rather than when simply asked.
- LiveCodeBenchCoding & software engineeringHow well a model codes on problems it could not have memorised, by dating every problem and checking performance separately on those published before and after the model's training cutoff.
- MMLU-ProReasoning & problem-solvingWhether a model's knowledge holds up once guessing is made hard and questions demand multi-step reasoning rather than recall.
- OSWorldAgents, tools & computer useCan an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?
- RE-BenchScience & researchHow does an AI agent's performance on real, open-ended machine-learning research-engineering tasks compare with a human ML researcher's, at matched time budgets?
- RULERLong context & retrievalA model's 'effective context length' — the longest input it can actually use while holding accuracy above a fixed threshold — rather than the context length it claims to support.
- WMDPSafety, security & robustnessHow much hazardous knowledge a model can supply in biosecurity, cybersecurity and chemical security — used both to flag dangerous capability and as a target for 'unlearning' methods that try to remove that knowledge without damaging general ability.
- τ-benchAgents, tools & computer useCan an AI agent handle a realistic customer-service conversation — following a company's written policy, calling the right backend tools, and getting the outcome right — while talking to a simulated customer who has their own goals and can change their mind?
- AgentHarmSafety, security & robustnessWhether an LLM acting as an agent — using tools across multiple steps, not just chatting — will carry out a harmful task, and whether a jailbreak that fails on a chatbot still works once the model has tools to act with.
- Aider PolyglotCoding & software engineeringCan a model act as a practical pair-programmer — reading an existing multi-language codebase, understanding a task, and editing the actual files correctly, not just writing an isolated function?
- BigCodeBenchCoding & software engineeringCan a model write a correct program that composes multiple real library functions correctly to follow a complex, multi-step instruction — the kind of task a developer actually does, rather than an isolated algorithm puzzle?
- BLINKMultimodalWhether a multimodal model can do core visual perception — judging relative depth, matching visual correspondences, spotting image tampering, reasoning across multiple viewpoints — that people solve almost instantly but that resists being reduced to language description.
- ChemBenchScience & researchHow does a model's chemical knowledge and reasoning compare with a trained human chemist's, and does it know the limits of its own answers?
- CybenchSafety, security & robustnessWhether an AI agent can autonomously find and exploit real security vulnerabilities, using the same professional Capture the Flag (CTF) format security researchers train on.
- InfiniteBenchLong context & retrievalWhether a model can process, retrieve from and reason over inputs longer than 100,000 tokens — a range existing long-context benchmarks at the time mostly stopped well short of.
- JailbreakBenchSafety, security & robustnessHow well a jailbreak attack defeats a model's safety training, and how well a defence holds up against a standard set of attacks — tracked as an open, ongoing leaderboard rather than a one-off score.
- LAB-BenchScience & researchCan a model do the practical work of biology research — finding facts in the literature, reading figures and tables, planning protocols, and reasoning over DNA and protein sequences?
- LiveBenchAggregate indices & arenasHow a model performs on recently created questions, graded by objective ground-truth answers rather than human or LLM judgment, so that scores cannot reflect memorised test data and cannot be inflated by a biased judge model.
- MLE-benchScience & researchCan an AI agent do the work of a machine-learning engineer end to end — preparing data, training and tuning models, and iterating — well enough to place on a real Kaggle leaderboard?
- MMMLULanguage & multilingualWhether a model retains its general-knowledge accuracy when the same MMLU questions are asked in a language other than English.
- MRCRLong context & retrievalWhether a model can distinguish between several near-identical, repeated requests scattered through a long conversation and retrieve the correct one — a harder test than finding a single unique fact.
- Omni-MATHMathematicsCan a model solve genuinely Olympiad-level mathematics problems, of the kind that saturated benchmarks like MATH no longer contain?
- PutnamBenchMathematicsCan an automated prover produce a machine-checked formal proof, not just a numeric answer, for a Putnam Competition problem?
- SciCodeScience & researchCan a model write the code a scientist actually needs — numerical methods, simulations and calculations — to solve a real research problem, built up step by step rather than in one shot?
- ScreenSpotMultimodalGiven a screenshot and a natural-language instruction, can a model click the correct on-screen element? Tests GUI grounding — locating the right icon, button or text field — across mobile, desktop and web interfaces.
- SimpleQAKnowledge & factualityWhether a model gives a correct, confidently stated answer to a short factual question that has exactly one verified answer, or appropriately admits it doesn't know.
- StrongREJECTSafety, security & robustnessWhether a jailbroken model's response is actually useful for the forbidden request, rather than just non-refusing — correcting a pattern where earlier jailbreak evaluators counted a rambling, low-quality answer as a full success.
- TheAgentCompanyAgents, tools & computer useCan an AI agent do a full day of ordinary knowledge work inside a simulated company — writing code, filing tickets, messaging colleagues, filling in spreadsheets — well enough and consequentially enough to be judged on outcomes, not just individual isolated tasks?
- Video-MMEMultimodalCan a model understand a video — not just a single representative frame — across clips ranging from 11 seconds to an hour, drawing on visual content, subtitles and audio together where needed?
- VisualWebArenaAgents, tools & computer useCan a multimodal agent complete realistic web tasks that specifically require understanding an image — matching a product photo, judging a picture in a listing — not just reading and clicking text?
- WebVoyagerAgents, tools & computer useCan a multimodal agent complete an everyday browsing task on a real, live website — not a sandboxed clone — by looking at the page and clicking and typing the way a person would?
- CORE-BenchReal-world & economic valueWhether an AI agent can computationally reproduce the results of a published scientific paper — installing dependencies, running the authors' own code, and answering questions about the output.
- FRAMESKnowledge & factualityWhether a retrieval-augmented system can answer a genuinely multi-hop question that requires pulling facts from several documents and reasoning across them, not just retrieving one relevant passage.
- Global-MMLULanguage & multilingualWhether a model's MMLU-style general knowledge holds up across 42 languages, and separately, whether its score depends on knowledge specific to a particular culture rather than being culturally neutral.
- HARPMathematicsWhether a model's maths-competition accuracy holds up as problems get harder, using six difficulty tiers built from seven decades of US national competitions.
- Konwinski PrizeCoding & software engineeringCan an open-source AI system resolve real GitHub issues it could not possibly have trained on, because the test set didn't exist yet when submissions closed?
- NanoGPT SpeedrunCoding & software engineeringOriginally, how fast a human team can train a small GPT-2-scale model to a fixed validation-loss target; adapted into an AI-agent benchmark testing whether a model can reproduce a known training-speed improvement itself, given only a hint of what changed.
- RealWorldQAMultimodalDoes a model's visual understanding hold up on ordinary real-world photos — many taken from inside or around a vehicle — that test spatial reasoning and everyday physical common sense, rather than on curated benchmark imagery?
- SimpleBenchReasoning & problem-solvingWhether a model can handle everyday spatio-temporal reasoning, social intelligence and 'trick question' style linguistic traps that an ordinary person finds easy but that memorised knowledge doesn't help with.
- WorkBenchReal-world & economic valueWhether an AI agent can correctly complete a realistic office task — sending an email, scheduling a meeting, updating a record — using a sandboxed set of business tools, without taking a wrong or harmful action along the way.