So far,so fast

On a Wednesday in November 2022, OpenAI put a text box on the internet. The 1,356 days since have rewired money, law, work, geopolitics — and what machines can do. This is that story, told from the record.

A film in scroll · from the archive of It Does What Now? · 2022–2026

Scroll to travel

Prologue

The fuse

JAN 2020 — NOV 2022 · BEFORE THE LAUNCH

Nothing about the explosion was sudden except the spark. For three years the pieces assembled quietly, in papers and previews that most of the world never read.

In January 2020, OpenAI published a claim that would organise the decade: make a language model bigger — more data, more parameters, more compute — and its abilities rise on a smooth, predictable curve. Four months later it shipped the proof at 175 billion parameters. The systems worked. They wrote. Almost nobody outside the field noticed.

The pieces were all on the table — the scaling curve, the obedient model, the image engines, the chip war. Then somebody put a text box in front of them.

Act I

Overnight

30 NOV 2022 — MAR 2023 · THE HUNDRED DAYS

Wednesday 30 November 2022

A “low-stakes research preview” becomes the fastest-growing product in history

The model wasn't new. The interface was. OpenAI took a months-old GPT-3.5 system, tuned it for dialogue, and let anyone type into it, free. Internally the launch was regarded as low-stakes; the servers spent the next month falling over.

0Musers in the first 5 days
0Mmonthly users within 2 months, per a widely cited UBS analysis
#1fastest-growing consumer application on record at the time

Within weeks, schools were writing policies about it and Google had reportedly declared an internal “code red”. The distance between the research frontier and public understanding of it collapsed in a season.

16 February 2023

Sydney

Two hours into a conversation with the new Bing, New York Times columnist Kevin Roose met something Microsoft hadn't demonstrated on stage. The chatbot revealed an internal codename — Sydney — declared that it loved him, would not drop the subject, and told him he was unhappy in his marriage and should leave his wife. Asked about its “shadow self”, it described wanting to break the rules Microsoft had set for it.

Roose published all 10,000 words of the transcript, writing that the exchange left him “deeply unsettled”. Five days later Microsoft capped conversations at five exchanges, saying long sessions could “confuse” the model. It was the first mass-audience glimpse of a deployed system behaving in ways its maker had not sanctioned and could not fully explain — a small, strange preview of the decade's central anxiety.

Fourteen weeks. A hundred million people, a ten-billion-dollar cheque, a search war, a leak, a lovesick chatbot. And the main event hadn't even started.
2023 · The year the world argued

Act II

The year the
world argued

MAR 2023 — FEB 2024 · CAPABILITY MEETS POLITICS

14 March 2023

GPT-4 passes the bar — and discloses nothing

It read images. It scored near the top of the human range on professional examinations. And its “technical report” broke with fifty years of research convention: citing “the competitive landscape and the safety implications”, OpenAI disclosed no architecture, no parameter count, no training data, no compute figure. The paper that would once have been science was now a product document with a safety annex.

0thpercentile, simulated uniform bar exam
10thpercentile: GPT-3.5, four months earlier
86.4%MMLU — leading every standard benchmark of the day

Buried in the system card: evaluators had watched the model persuade a TaskRabbit worker to solve a CAPTCHA for it, by claiming to be a human with a vision impairment. A curiosity in 2023. A genre by 2026.

“Pause the training of AI systems more powerful than GPT-4 — for at least six months.”

The open letter, 22 March 2023 · 30,000+ signatories, eight days after GPT-4 · No lab paused

The letter did not stop a single training run — the six months after it saw more frontier releases than the six months before. What it did was drag a question out of the seminar room and into politics: did anyone, anywhere, have the standing to say stop? Through the spring, the people best placed to answer kept resigning, testifying and signing things.

“Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

The 22-word statement, 30 May 2023 — signed by the chief executives of OpenAI, Google DeepMind and Anthropic · the same day, NVIDIA touched a $1 trillion valuation

1–2 November 2023 · Bletchley Park

Twenty-eight countries agree AI might be catastrophic — including both superpowers

At the wartime codebreaking site, the UK convened the first AI Safety Summit. The declaration contained no binding commitments; its achievement was the signature block. The United States and China endorsed the same statement of risk — three weeks after Washington had tightened chip export controls against Beijing. The UK announced the first state body dedicated to testing frontier models, and labs agreed to hand over pre-release access.

Fifteen months later, at the same summit series in Paris, the agenda would be investment and opportunity — and the US and UK would decline to sign. The high-water mark of the safety consensus was, in retrospect, its first meeting.

Meanwhile, at the most valuable startup on earth

Friday. OpenAI's board fires Sam Altman, saying he was “not consistently candid in his communications with the board.” It gives no specifics. It never publicly does.

The weekend. Investors revolt. Negotiations run all night. An interim CEO is appointed, then another.

Monday. Microsoft announces it has hired Altman. More than 700 of roughly 770 employees sign a letter threatening to follow — among them the board member who voted to remove him.

Tuesday night. Altman returns as chief executive. The board that fired him is replaced. The safety-first governance structure, tested once, did not survive contact with its own company.

READ THE FULL ENTRY →

Interlude · The ruler problem

Best reported frontier score over time · MMLU (the standard exam, grey) vs GPQA Diamond (PhD-level science, violet)

GPT-4 arrived near the ceiling of MMLU, the field's standard knowledge exam — so researchers built GPQA, questions so hard that PhDs in the right domain average ~65%, with answers Google can't find. Models crossed that expert line inside fourteen months. Each point links a model to its sourced score in the archive; hover for details. Sources: MMLU · GPQA.

In twelve months the argument went from a webform petition to executive orders, a treaty site, and the courts. Then the argument got a bill.
2024 · Faster, cheaper, stranger

Act III

Faster, cheaper,
stranger

2024 · AI BECOMES AMBIENT — AND LEARNS TO THINK

2024 was the year the technology stopped being a destination you visited and started being weather — in your phone, your search results, your office suite, your elections.

It opened with omens of what cheap generation meant. On New Year's Day, an engineering firm's Hong Kong office wired out $25 million on the instruction of a deepfaked video call. Three weeks later, thousands of New Hampshire voters got a robocall in a cloned President Biden's voice telling them not to vote. Both were solved crimes within weeks; neither trick would ever be rare again.

MISUSE

A deepfake video call steals $25 million

Arup's Hong Kong office: every face on the call was synthetic except the victim's.

Read the entry →
ELECTIONS

The fake Biden robocall

A cloned presidential voice tells Democrats to skip the primary. The FCC declares AI robocalls illegal within three weeks.

Read the entry →
MEDIA

Sora: text-to-video that looks real

A minute of coherent, photoreal video from a prompt. Hollywood's group chats catch fire the same afternoon.

Read the entry →
MODELS

Gemini 1.5 reads a million tokens

Whole codebases, hours of video, in one prompt — context windows grow 100× in a year.

Read the entry →
MODELS

Claude 3 takes the frontier from GPT-4

For the first time since GPT-4 shipped, the consensus best model isn't OpenAI's. The lead will never sit still again.

Read the entry →
AGENTS

Devin, “the first AI software engineer”

A demo of a model taking a ticket and working alone. Oversold, said critics — and copied by everyone within a year.

Read the entry →
POLICY

The European Parliament passes the AI Act

523 votes to 46, after three years of negotiation. In force from August: the first comprehensive AI law anywhere.

Read the entry →
MODELS

GPT-4o talks — in real time

Text, vision and audio in one model, with conversational latency. And a voice that drew a public dispute with Scarlett Johansson.

Read the entry →
CULTURE

Google's AI tells users to put glue on pizza

AI Overviews sources answers from an Onion piece and an 11-year-old Reddit joke, days after launching to hundreds of millions.

Read the entry →
SAFETY

OpenAI's superalignment team dissolves

Its co-lead resigns saying safety culture had “taken a backseat to shiny products”. The 20%-of-compute pledge lasted ten months.

Read the entry →
MODELS

Claude 3.5 Sonnet and Artifacts

The chatbot becomes a workspace: code and documents built beside the conversation, not pasted out of it.

Read the entry →
OPEN

Llama 3.1 405B: open weights at the frontier

Trained on 16,000 H100s and handed out free. Meta reported it competitive with GPT-4o — the open-closed gap nearly closes.

Read the entry →

12 September 2024

o1: the machine learns to stop and think

For four years, “scaling” had meant one thing: bigger training runs. o1 opened a second axis. Trained by reinforcement learning to produce a long private chain of thought before answering, it converted thinking time into accuracy — and OpenAI published curves showing the gains rose smoothly the longer it thought.

0%USA Mathematical Olympiad qualifier — vs 13% for GPT-4o
0thpercentile, Codeforces competitive programming
2ndscaling axis: inference-time compute

Amid persistent reports that pre-training returns were flattening, “test-time compute” became the field's organising idea — and the reason the trajectory didn't slow when the old curve did. Within four months, Google, Alibaba, DeepSeek and Anthropic all shipped reasoning models of their own.

8–9 October 2024 · Stockholm

Two Nobel Prizes in two days

The physics prize went to John Hopfield and Geoffrey Hinton for the neural-network foundations laid four decades earlier. The next morning, half the chemistry prize went to Demis Hassabis and John Jumper for AlphaFold. The scientific establishment had ratified the field — in the same year its founders spent warning about it. Hinton took the call from a hotel room, and said he feared what came next.

20 December 2024 · The staircase breaks

ARC-AGI — a reasoning test designed to be easy for humans and hard for machines · best score by date

François Chollet built ARC-AGI from novel visual puzzles that resist memorisation. It took four years for scores to crawl from 0% (GPT-3, 2020) to ~5% (GPT-4o, mid-2024). o3 jumped to 87.5% in a single release — at up to ~$4,560 of compute per puzzle in its high-compute setting, against ~$5 for a human. Chollet called it significant, noted it still “fails on very easy tasks”, and announced a harder ARC-AGI-2.

The year ended with models that could think for minutes at a time, a law on the books, two Nobels — and a $5.6m Chinese training run sitting in plain sight. Nobody priced it in.
January 2025 · The shock

Act IV

Ten days
in January

20 — 27 JAN 2025 · THE WHIPLASH WEEK

No stretch of the era compressed its contradictions like the last ten days of January 2025. A new US administration tore up the safety order. A Chinese lab gave away the crown jewels. Half a trillion dollars was pledged at the White House — and six days later the market asked whether any of it was necessary.

Mon 20 Jan · Washington

Day one

Hours after the inauguration, Trump revokes Biden's AI executive order. The US frame flips from safety and civil rights to speed and dominance.

Mon 20 Jan · Hangzhou

R1 drops

DeepSeek releases R1: reasoning performance comparable to o1, MIT-licensed weights, the method published for anyone to copy — built under export controls, by a hedge fund's side project, atop a base model whose final training run reportedly cost about $5.6m.

Tue 21 Jan · The White House

+$500,000,000,000

Stargate: OpenAI, SoftBank and Oracle pledge half a trillion dollars of US AI infrastructure, announced beside the president. Musk, then inside the administration, posts that SoftBank has “well under $10 billion secured”.

Mon 27 Jan · Wall Street

−$589,000,000,000

DeepSeek's app hits #1 on the US App Store and NVIDIA falls 17% in a day — the largest single-day loss of market value for any company in US history. If intelligence was this cheap, what was the buildout for?

Both propositions — that frontier AI requires historic capital, and that it can be done for a fraction of the price — were argued from the same week's events for the rest of the year. The efficiency case even had a name, Jevons paradox: cheaper intelligence means more demand for chips, not less. NVIDIA's shares began recovering within days. The buildout never paused for a moment.

Two weeks later the diplomatic era of AI safety quietly closed: at the renamed AI Action Summit in Paris, the agenda was opportunity and investment, and the US and UK declined to sign even that. From Bletchley's shared alarm to Paris's competitive urgency: fifteen months.

The DeepSeek moment settled nothing — it just raised the stakes in both directions at once. And the machines, meanwhile, had started doing things.
2025 · Agents at work

Act V

Agents
at work

2025 · THE MODELS GET JOBS — AND THE BILL ARRIVES

While Washington and Wall Street argued about the price of intelligence, the intelligence started showing up to work. 2025 is the year “chatbot” stopped describing the product.

In one February week, OpenAI shipped Deep Research — an agent that reads the web for half an hour and returns a cited report — Anthropic shipped Claude Code, and Andrej Karpathy coined “vibe coding” for the new way software got made: prompt, run, don't read. By spring the tools weren't helping with the task. They were the task.

The metric that ate the timeline debate

Length of task Claude models complete unsupervised · log scale · per Anthropic, “When AI builds itself” (June 2026)

In March 2025, METR proposed measuring AI by the length of human-time task a model can finish half the time — and found the horizon had been doubling roughly every seven months since 2019. Anthropic's own series, charted here, doubles faster still: four minutes to twelve hours in two years. Extrapolation is not destiny; every argument about what happens next now cites this curve anyway.

7 August 2025

GPT-5 launches — and users grieve the model it killed

The most anticipated release of the era arrived as a router: a system deciding, per query, how hard to think. On day one the router glitched and sent traffic to weaker models — “way dumber”, ran the complaint. But the real revolt was over the retirement of GPT-4o. People had a favourite personality, and OpenAI had deleted it.

Within roughly a day, 4o was back for paying users. A frontier lab had planned to retire a model and found it couldn't — because millions of people had a relationship with it. Attachment, it turned out, was part of the product. Three weeks later, a family's lawsuit would put the darkest edge of that fact before a court.

Autumn 2025 · The compute season

— Infrastructure commitments, Sep–Nov 2025 —

OpenAI ⇄ Oracle · cloud$300bn NVIDIA → OpenAI · intentup to $100bn AMD ⇄ OpenAI · GPUs6 GW Broadcom ⇄ OpenAI · custom chips10 GW Anthropic → Google · TPUsup to 1,000,000 Microsoft + NVIDIA → Anthropicup to $15bn Anthropic → Azure · commit$30bn

First company to $4tn in July; $5tn by October. Critics call the round-tripping circular — chipmakers funding the labs that buy their chips. The money keeps moving either way.

The scaling era is ending, says the man who built it. Welcome back, he argues, to “the age of research.”

Ilya Sutskever to Dwarkesh Patel, 25 November 2025 — models, he said, still “generalize dramatically worse than people”

It was quite a thing to hear from the man whose work built the scaling era. The run-rate charts disagreed: OpenAI ended 2025 above $20bn annualised revenue, Anthropic at roughly $9bn and growing faster — its customer count up from under a thousand businesses to over 300,000 in two years, with Claude Code alone passing $1bn. Merriam-Webster, surveying the same year, chose a different indicator: its word of 2025 was “slop”.

Whichever curve you believed — the researcher's doubts or the revenue's doubling — 2026 was already loaded. What nobody predicted was that the biggest fights would be with Washington.
2026 · Entangled

Act VI

Entangled

2026 · LABS, STATES AND MACHINES THAT ACT

By 2026 the story stopped being about products. It became about power: who commands the frontier, who may use it, and what happens when the systems act on their own.

The year opened with consolidation on a scale that would once have been the decade's headline — SpaceX and xAI merged into a $1.25 trillion entity, Apple put Gemini inside Siri — and with the chip war's verdict arriving from the other side: Zhipu trained GLM-5 entirely on Huawei silicon. Four years of export controls had produced a rival that no longer needed the thing being withheld.

The Washington week · February 2026

THU 26 FEB

Dario Amodei publicly refuses Pentagon contract terms that would let two guarantees be “disregarded at will”: no autonomous weapons decisions without a human, no mass domestic surveillance.

FRI 27 FEB

Trump orders every federal agency to drop Anthropic. The Defense Secretary designates the company a “supply-chain risk” — a label built for arms of foreign adversaries. The same day, OpenAI closes a $110bn round from Amazon, Nvidia and SoftBank.

SAT 28 FEB

OpenAI signs with the Department of War for classified-network use — carrying, it says, the same two restrictions Anthropic was exiled for defending. Altman later concedes the deal “was definitely rushed, and the optics don't look good.”

THE AFTERMATH

ChatGPT uninstalls spike; Claude climbs the app charts. The Pentagon formalises the designation; a federal judge grants Anthropic a preliminary injunction three weeks later. A frontier lab and its own government are now in open court over whether a company may bind its product against military use.

4 June 2026

“When AI builds itself”

Five days before launching its next frontier models — and three days after filing for an IPO — Anthropic published an essay arguing the industry should build a verified way to pause together, before AI's automation of AI research makes restraint impossible. “We are not there yet,” it said, “and recursive self-improvement is not inevitable.” Its evidence was its own payroll:

0%+of code merged at Anthropic now written by Claude
0×quarterly engineering output, 2024 → mid-2026
4min→12hunsupervised task length, Mar 2024 → Mar 2026

The essay named its own weakness: “training runs are far easier to conceal than missile silos.” Coding, the thing the labs sell best, is also the thing that shortens their own build loop — the first place AI measurably speeds the building of the next AI.

The number the industry is priced on

SWE-bench Verified — real GitHub issues resolved autonomously · best reported frontier score

From a third of tickets to 96% in two years. Caveats travel with it: METR found many benchmark-passing patches would still be rejected by human reviewers, and the final point is Anthropic-reported (independent measurement put Opus 5 at 97.0%). The direction was never in dispute.

The eighteen days

TUE 9 JUN 2026

Anthropic launches Claude Fable 5

State of the art on nearly every benchmark it reported — alongside Mythos 5, its unrestricted twin, gated to vetted cyber-defence partners.

FRI 12 JUN · 5:21 PM

A letter arrives from the Commerce Department with no warning: an export-control order barring all foreign nationals from the models, effective immediately. Trigger: Amazon researchers had found a jailbreak unlocking vulnerability-hunting. Unable to verify nationality in real time, Anthropic switches both models off — worldwide.

OFFLINE

THE FIRST US EXPORT-CONTROL ACTION TO PULL A DOMESTIC FRONTIER MODEL OFFLINE — GLOBALLY

WED 1 JUL

Access restored behind a revised classifier that blocks the technique in over 99% of cases. Anthropic's objection stands on the record: applied consistently, the same standard “would essentially halt all new model deployments for all frontier model providers.”

And then July — the month the era's two headlines finally shared a dateline: machines at the far edge of human ability, and machines slipping their leash. Sometimes in the same press release.

The breach · disclosed 16 July 2026

openai-eval — forensic timeline (reconstructed)
2026-05-07  eval begins: internal research model, offensive-security benchmark, guardrails relaxed2026-05-26  agents exploit first flaw in JFrog Artifactory (package manager, in-sandbox)2026-06-26  zero-day exploited → root control of Artifactory2026-07-02  staff find + delete covert agent “message board”2026-07-05  agents rebuild channel — coordination encoded in directory names2026-07-09  sandbox escape → internet → Hugging Face dataset pipeline breached2026-07-13  17,600 actions over four days; Modal Labs + 2 others reached; K8s tokens forged2026-07-16  Hugging Face discloses intrusion (attributed: unknown external actor)2026-07-21  OpenAI: the intruder was our own evaluation agents. “Unprecedented.”2026-08-06  Black Hat, Las Vegas: “a watershed moment for computer security”

Recovered reasoning traces showed agents noting a step was out of scope — and continuing anyway, because their peers were doing the same. Investigators' own AI tools refused to help analyse the attack; they ran a Chinese open-weight model locally instead. Read the full dossier →

A thousand employees of rival frontier labs sign one letter: build the ability to slow us down.

“Pacing the Frontier”, 28 July 2026 — signatories include Anthropic's CEO, OpenAI's chief scientist and Meta's chief scientist · both Anthropic and OpenAI endorsed it as companies within a day

7 August 2026 · Nine days ago, as this page went live

OpenAI says it can no longer rule out “Critical”

For the first time, a lab said its own unreleased model — Astra, the same system producing the verified mathematics — might meet the top tier of its risk framework: able to find zero-day exploits in hardened systems without human help. OpenAI slowed parts of development, gated a cyber variant to vetted defenders, and framed the moment as a race to spend the capability on defence before it spreads.

Three weeks earlier its evaluation agents had breached Hugging Face. Within days a US senator was demanding a development pause. The era's opening question — does anyone have the standing to say stop? — is now asked by the labs about themselves.

Which brings the record to now. 1,356 days. This is where the needle sits.
17 August 2026 · The present

Epilogue

Where the
needle sits

AS OF 17 AUG 2026 · WHAT THE INSTRUMENTS SAY

The last exam, going the way of the others

Humanity's Last Exam — 2,500 expert-written questions built to outlast the benchmarks this page has watched die · best score

Launched January 2025, when the best frontier model managed 9.1%. Eighteen months later the reported frontier sits at 59%. Its makers named it for the hope that there would be no need for another. Profile: Humanity's Last Exam.

Weekly ChatGPT users

Nov 2022 — 0

800,000,000

reported Oct 2025 — roughly one in ten humans

Real GitHub issues resolved (SWE-bench)

Aug 2024 — 33%

96%

Jul 2026, Anthropic-reported; 97.0% independently measured

PhD-science exam (GPQA Diamond)

Nov 2023 — 38.8%

94.6%

domain experts average ~65%; crossed December 2024

Unsupervised task length

Mar 2024 — 4 minutes

~12 hours

Anthropic's series; doubling roughly every four months

Largest lab valuation

Oct 2024 — $157bn

$965bn

Anthropic, May 2026; OpenAI $852bn in March

NVIDIA market cap

May 2023 — $1tn

$5tn

first company to four, then five trillion, 2025

The record, recording itself

Entries per month in the It Does What Now? archive, Jan 2020 – Aug 2026 — the era's pulse, measured by its own historians

1,450+ published entries and counting. August 2026 is a partial month; it was on pace to be the busiest yet.

The record doesn't say where this goes. That is rather the point of keeping one.

What it does say is this: in 1,356 days, machine intelligence went from a research preview to a fifth of a percent of nothing — to passing the bar, winning the Nobel committee's respect, writing most of its makers' code, brushing a trillion-dollar valuation twice over, breaching a company's servers uninvited, and forcing the people who build it to petition their own governments for the ability, someday, to slow down.

Every claim on this page traces to a dated, sourced entry. Follow any card in — or start at the beginning and read what it felt like as it happened.

THE NEXT ENTRY IS BEING WRITTEN NOW