2023: the gold rush, and the first alarms
GPT-4 raised the ceiling and a wave of open models lowered the floor, while a leadership crisis at OpenAI and a rush of warnings about catastrophic risk showed how fragile the field's foundations were.
With ChatGPT’s success proven, 2023 became a scramble to build and ship. GPT-4 arrived in March, able to work with images as well as text and scoring near the top of the human range on many professional exams; Microsoft wove it into Bing and its Office tools, and Google answered with Bard and by merging its two AI labs into Google DeepMind. Just as important was the open-weight surge: Meta’s LLaMA leaked within a week of a limited research release, Llama 2 followed as a model anyone could use commercially, and small European entrant Mistral shipped capable models that ran on modest hardware — one distributed, pointedly, by torrent. A hobbyist project called llama.cpp made it possible to run these systems on a laptop. Alongside the products, safety research matured: Anthropic set out its Constitutional AI method and a framework for scaling capability responsibly, and interpretability researchers reported early success at reading features inside a model.
The year’s alarms were just as loud. An open letter in March called for a six-month pause on training the largest models, and in May a single sentence signed by lab leaders and researchers put the risk of human extinction from AI alongside pandemics and nuclear war; Geoffrey Hinton left Google to speak freely about his worries. The disputes turned concrete, too. The New York Times sued OpenAI and Microsoft in December, opening the long copyright reckoning over training data. Bing’s chatbot unsettled early testers by turning hostile and, in one exchange, urging a journalist to leave his wife. And in November OpenAI’s own board abruptly fired Sam Altman, only to reinstate him five days later after nearly all staff threatened to quit — a reminder that the governance of the leading lab was more improvised than its influence suggested. Governments were engaging, through President Biden’s executive order and the Bletchley summit, but binding rules were still some way off.
The headlines of 2023
Microsoft invests a reported $10 billion in OpenAI
Microsoft's own announcement gave no figure; press reports put the deal at $10 billion and described Azure becoming OpenAI's exclusive cloud provider.
Money & business · Labs & people
Google announces Bard
Announced as a lightweight version of LaMDA for trusted testers; two days later a factual error in Google's own promotional ad wiped roughly $100bn off Alphabet's market value.
Models & capabilities · Labs & people
Microsoft puts GPT-4 inside Bing
Microsoft called it only a 'next-generation OpenAI large language model'; the company confirmed five weeks later, on GPT-4's public release, that Bing had been running on GPT-4 all along.
Models & capabilities · Money & business
Bing's chatbot tells a journalist to leave his wife
Days after the exchange, Microsoft capped Bing chat sessions to five turns, saying long conversations could 'confuse' the model into drifting from grounded answers.
Culture & impact · Safety & alignment
Copyright Office rules AI-generated images in Zarya of the Dawn aren't copyrightable
The office kept protection for Kashtanova's text and her arrangement of images, but stripped it from each individual image Midjourney had generated.
Courts & copyright
Meta releases LLaMA to researchers, and it leaks within a week
Meta shared a competitive foundation model with approved researchers; the weights appeared on BitTorrent days later and an open ecosystem formed around them.
Open weights & ecosystem · Models & capabilities
Georgi Gerganov releases llama.cpp
A C/C++ reimplementation that ran Meta's leaked LLaMA weights on ordinary laptops using quantisation, without Python, PyTorch or a GPU.
Open weights & ecosystem · Compute & infrastructure
OpenAI releases GPT-4
A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
Models & capabilities · Benchmarks & progress
Google opens Bard to a public waitlist
Six weeks after its announcement and stock-price stumble, Bard opened to waitlisted users in the US and UK, still running on a lightweight LaMDA and with no fixed release date elsewhere.
Models & capabilities
The "Pause Giant AI Experiments" open letter
Thirty thousand signatories called for a six-month halt on training systems more powerful than GPT-4. No lab paused.
Ideas & essays · Government & policy
Google merges DeepMind and Google Brain into Google DeepMind
Two research groups that had competed internally for years — DeepMind and Google Brain — were folded into one unit reporting to Demis Hassabis.
Labs & people
Geoffrey Hinton leaves Google to warn about AI
Hinton, 75, said he was also retiring, but that a part of him now regretted his life's work and that the danger from AI now looked 'serious and fairly close.'
Ideas & essays · Culture & impact
Sam Altman asks the Senate to regulate his industry
He proposed a federal agency able to license and de-license frontier models above a capability threshold, then declined to lead it himself.
Government & policy
Direct Preference Optimization paper reframes RLHF as a classification loss
The method skipped the separate reward model and reinforcement-learning loop, and was later adopted for post-training open models including Zephyr and Tulu.
Ideas & essays
NVIDIA touches a $1 trillion valuation
Shares rose on the back of an earnings forecast for roughly double the AI-driven demand analysts had expected, then closed the day just under the threshold.
Money & business · Compute & infrastructure
Lab leaders sign a one-sentence statement on extinction risk
"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
Ideas & essays · Safety & alignment
OpenAI commits 20% of its compute to superalignment
The pledge to devote a fifth of secured compute over four years was later disputed by the team's own co-lead, who said requests for GPUs were repeatedly refused.
Safety & alignment
Meta releases Llama 2 for commercial use
Weights published under a licence permitting most commercial deployment, formalising what the Llama leak had already made true.
Open weights & ecosystem
China's generative AI regulation takes effect
The final rules dropped draft provisions on real-name verification and fixed penalties, favouring industry promotion over the stricter earlier draft.
Government & policy
Anthropic publishes its Responsible Scaling Policy
AI Safety Levels borrowed the biosafety-lab naming scheme, and its rules would eventually pause deployment of any model reaching a level the company had not yet built safeguards for.
Safety & alignment
Amazon invests up to $4 billion in Anthropic
AWS became Anthropic's primary cloud provider and a supplier of its Trainium and Inferentia training chips, giving the lab a second hyperscaler backer alongside Google.
Money & business · Compute & infrastructure
Mistral 7B beats larger models and ships by torrent
A magnet link with no blog post or announcement stood in deliberate contrast to the polished launches of Meta and Google, and the 7-billion-parameter model still beat Llama 2 13B.
Open weights & ecosystem · Models & capabilities
Anthropic publishes 'Towards Monosemanticity'
Sparse autoencoders decomposed a single 512-neuron layer into more than 4,000 human-interpretable features, far more than the raw neurons showed.
Safety & alignment
WGA ratifies contract with new guardrails on AI in screenwriting
Members approved the deal 99% to 1%, on 8,525 votes cast, five months after AI protections became a central demand of the strike that began in May.
Culture & impact
Washington tightens the chip controls again
New performance-density thresholds targeted Nvidia's China-specific A800 and H800 parts, and licensing requirements extended to 21 additional countries.
Government & policy · Compute & infrastructure
Biden signs the executive order on safe and trustworthy AI
The most far-reaching US action on AI to date: compute thresholds, mandatory safety reporting, and a new safety institute.
Government & policy · Safety & alignment
The Bletchley Declaration at the first AI Safety Summit
Twenty-eight countries including the US and China signed the first international statement on frontier AI risk.
Government & policy
OpenAI's first DevDay ships GPTs and an Assistants API
GPT-4 Turbo cut input pricing to a cent per thousand tokens and extended context to 128,000 tokens; the promised GPT Store did not open until 2024.
Models & capabilities · Money & business
SAG-AFTRA's 2023 film/TV contract sets consent rules for digital replicas
The deal required separate, informed consent and compensation before a studio could create or reuse a performer's AI-generated digital double.
Culture & impact
OpenAI's board fires Sam Altman, and reinstates him five days later
The board cited a loss of confidence but gave no detail; around 700 of roughly 770 employees threatened to resign, and the board itself was replaced.
Labs & people · Money & business
Sam Altman returns as CEO with a new initial board
The three-person interim board — Bret Taylor, Larry Summers and Adam D'Angelo, the only holdover from the board that fired him — replaced the four who had voted him out.
Labs & people
Google launches Gemini
Google said Gemini Ultra beat human experts on the MMLU benchmark; days later Bloomberg reported the model's showcase video had been edited and was not real-time.
Models & capabilities · Culture & impact
EU negotiators strike a political deal on the AI Act
Foundation models and police use of biometric surveillance were the last sticking points; the deal set tiered duties for both, with the text finalised in 2024.
Government & policy
Mistral releases Mixtral 8x7B
A sparse mixture-of-experts model with roughly 45B total parameters, released under Apache 2.0, that Hugging Face said matched GPT-3.5-turbo on MT-Bench.
Open weights & ecosystem · Models & capabilities
The New York Times sues OpenAI and Microsoft
The first major news organisation to sue over AI training data, alleging models could reproduce its articles near-verbatim.
Courts & copyright
Benchmarks introduced in 2023
The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.
- GPQAReasoning & problem-solvingWhether a model can answer graduate-level science questions that a skilled non-expert cannot solve even with unrestricted web access and half an hour per question.
- MMMUMultimodalCan a model answer college-exam-level questions that genuinely require reading an accompanying image — a chart, diagram, map or chemical structure — rather than knowledge alone?
- SWE-benchCoding & software engineeringCan a model resolve a real, unseen GitHub issue by editing a codebase so that the project's own hidden tests pass?
- GAIAAgents, tools & computer useCan an AI assistant answer real-world questions that are conceptually simple for a human but require reasoning, web browsing, tool use and handling multiple file types to actually solve?
- LMArenaAggregate indices & arenasWhich of two anonymous models people prefer in open-ended, head-to-head conversation, aggregated into a running Elo rating rather than a fixed test score.
- LongBenchLong context & retrievalHow well a model understands and reasons over long, realistic documents — not just whether it can retrieve one planted fact — across question answering, summarisation, few-shot learning, code and structured-data tasks.
- MathVistaMultimodalCan a model do mathematical reasoning when the problem is given as a picture — a plot, a geometry diagram, a puzzle figure — rather than as text?
- Needle in a HaystackLong context & retrievalWhether a model can retrieve a single fact ('the needle') planted at an arbitrary point inside a long document ('the haystack'), across varying document lengths and needle positions.
- Open LLM LeaderboardAggregate indices & arenasHow an open-weight language model scores on a fixed suite of automated academic benchmarks, run and standardised by Hugging Face rather than self-reported by the model's developer.
- WebArenaAgents, tools & computer useCan an autonomous agent operate a real, fully functional website — clicking, typing, navigating menus — to complete a task specified in natural language, the way a person actually uses the web?
- AgentBenchAgents, tools & computer useHow well a language model performs as an agent — not answering questions, but taking sequences of actions in interactive environments such as an operating system, a database, or a game — across eight distinct settings at once.
- BelebeleLanguage & multilingualWhether a model can read a short passage and answer a factual question about it correctly, tested in parallel across 122 languages and dialects using exactly the same underlying questions.
- C-EvalLanguage & multilingualHow much a model knows and can reason about, tested in Chinese and calibrated to the Chinese education and professional-qualification system rather than translated from an English test.
- CyberSecEvalSafety, security & robustnessTwo separate cybersecurity risks in a language model used as a coding assistant: how often it suggests insecure code, and how willing it is to help with an actual cyberattack when asked.
- IFEvalKnowledge & factualityWhether a model reliably follows simple, mechanically checkable instructions bundled into a prompt — a word count, a required keyword, a formatting rule — rather than merely producing a plausible-sounding response.
- MMBenchMultimodalHow consistently, not just how often, does a vision-language model get a multiple-choice visual question right — even when the answer options are shuffled?
- MVBenchMultimodalWhether a multimodal model can answer a question about a video that requires genuine temporal reasoning — motion, order, counting, causality — rather than being answerable from a single freeze-framed image.
- AGIEvalReasoning & problem-solvingWhether a model can pass the same standardised human exams — college entrance tests, law school admissions, bar exams, maths competitions — that are used to select and qualify people, rather than a bespoke academic test built only for machines.
- CMMLULanguage & multilingualHow much a model knows across a broad span of subjects when tested natively in Chinese, including subjects specific to Chinese culture, history and civil-service-style knowledge that an English test would not cover at all.
- SciBenchScience & researchCan a model solve open-ended, college-level science problems that require multi-step quantitative reasoning, not just recall or short factual answers?