The chatbot arms race
The public contest to ship the leading consumer chatbot — ignited by ChatGPT, and since then a frontier that changes hands between OpenAI, Google and Anthropic almost every quarter.
The chatbot arms race is the public, product-facing contest to ship the most capable conversational model — the visible surface of a deeper competition over frontier capability. Its groundwork was GPT-3 and the InstructGPT fine-tuning that made a raw model follow instructions, but the race proper began with ChatGPT, which reached a hundred million users faster than any consumer product before it.
The incumbents scrambled. Google announced Bard and Microsoft put GPT-4 inside Bing within a day of each other, the latter promptly unsettling a journalist in extended conversation. GPT-4 reset the bar; Google answered by merging its labs and launching Gemini.
From 2024 the frontier changed hands quarter by quarter. Claude 3 briefly took the lead from GPT-4; GPT-4o added real-time voice; and Gemini, Claude and OpenAI traded the top of the reasoning benchmarks through 2025. GPT-5 launched to a backlash over the model it retired, before Gemini 3 and then Claude Opus 5 pushed ahead again. The thread’s defining feature is its churn: no lead in it has lasted a year.
Facebook releases the Blender chatbot
Released in three sizes up to 9.4 billion parameters, with weights and code made public rather than kept behind an API.
Models & capabilities · Open weights & ecosystem
OpenAI publishes the GPT-3 paper
A 175-billion-parameter language model that learned new tasks from examples in its prompt, without any weight updates.
Models & capabilities · Ideas & essays
GPT-3 demos spread beyond the research community
Developer Sharif Shameem's demo turning plain-English descriptions into working webpage code was among the clips that drew attention beyond NLP researchers.
Culture & impact
The Guardian publishes an op-ed written by GPT-3
Editors ran the model eight times on the same prompt and spliced the best passages together, a process disclosed in a footnote that critics said undercut the framing.
Culture & impact
Google announces LaMDA and TPU v4 at I/O
A single TPU v4 pod combined 4,096 chips for over one exaflop of compute, while LaMDA was pitched on open-ended conversation rather than benchmark scores.
Models & capabilities · Compute & infrastructure
AI21 Labs releases Jurassic-1
Jurassic-1 Jumbo's 178 billion parameters slightly exceeded GPT-3's, and its 250,000-token vocabulary — five times GPT-3's — aimed to cut per-word token costs.
Models & capabilities
OpenAI removes the GPT-3 waitlist
General availability in supported countries turned the API from a curated experiment into a commodity developers could just buy, though several countries were excluded.
Models & capabilities · Money & business
OpenAI ships InstructGPT and makes RLHF the default
Models fine-tuned on human preference data were preferred to a model a hundred times larger, reframing alignment as a product feature.
Models & capabilities · Safety & alignment · Ideas & essays
OpenAI launches ChatGPT
A free web interface to an existing GPT-3.5 model, reported by one analysis to have reached 100 million monthly users within two months.
Models & capabilities · Culture & impact
Google is reported to declare a code red over ChatGPT
Teams from research and Trust and Safety were reportedly reassigned to accelerate AI products, with an internal target tied to Google's May developer conference.
Labs & people
Google announces Bard
Announced as a lightweight version of LaMDA for trusted testers; two days later a factual error in Google's own promotional ad wiped roughly $100bn off Alphabet's market value.
Models & capabilities · Labs & people
Microsoft puts GPT-4 inside Bing
Microsoft called it only a 'next-generation OpenAI large language model'; the company confirmed five weeks later, on GPT-4's public release, that Bing had been running on GPT-4 all along.
Models & capabilities · Money & business
Bing's chatbot tells a journalist to leave his wife
Days after the exchange, Microsoft capped Bing chat sessions to five turns, saying long conversations could 'confuse' the model into drifting from grounded answers.
Culture & impact · Safety & alignment
Meta releases LLaMA to researchers, and it leaks within a week
Meta shared a competitive foundation model with approved researchers; the weights appeared on BitTorrent days later and an open ecosystem formed around them.
Open weights & ecosystem · Models & capabilities
OpenAI opens the ChatGPT API at a tenth of the price
gpt-3.5-turbo was priced at $0.002 per 1,000 tokens, roughly a tenth of the previous GPT-3.5 rate, alongside a new Whisper transcription API.
Money & business · Models & capabilities
Inflection AI launches out of stealth
The company, structured as a public benefit corporation, announced Pi, a conversational assistant it described as a supportive companion rather than a search or productivity tool.
Labs & people
Anthropic launches Claude
The company's first public assistant launched, via a chat interface and API, on the same day OpenAI released GPT-4.
Models & capabilities · Labs & people
OpenAI releases GPT-4
A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
Models & capabilities · Benchmarks & progress
Google opens Bard to a public waitlist
Six weeks after its announcement and stock-price stumble, Bard opened to waitlisted users in the US and UK, still running on a lightweight LaMDA and with no fixed release date elsewhere.
Models & capabilities
Alibaba launches Tongyi Qianwen chatbot
Launched without advance notice and restricted to corporate clients and select media on an invite-only basis; Alibaba did not disclose a parameter count.
Models & capabilities
Google merges DeepMind and Google Brain into Google DeepMind
Two research groups that had competed internally for years — DeepMind and Google Brain — were folded into one unit reporting to Demis Hassabis.
Labs & people
Inflection AI launches Pi, a personal AI chatbot
The company, co-founded by DeepMind's Mustafa Suleyman and LinkedIn's Reid Hoffman, positioned Pi as a supportive companion rather than a productivity tool.
Models & capabilities
iFlytek launches Spark (Xinghuo) cognitive model
Chairman Liu Qingfeng said the model beat ChatGPT on Chinese-language tests and pledged to match it in English by October 2023, at a Hefei launch event.
Models & capabilities
Google answers with PaLM 2 and a global Bard
Bard moved onto the new model, dropped its waitlist, and expanded to over 180 countries, adding Japanese and Korean with 40 languages planned.
Models & capabilities
Anthropic releases Claude 2
A 100,000-token context window and a jump to 71.2% on the Codex HumanEval coding test, alongside a consumer web app opened to the US and UK.
Models & capabilities
OpenAI launches ChatGPT Enterprise
The tier offered unlimited GPT-4 access at double speed, a 32,000-token context window and a pledge not to train on business data.
Models & capabilities · Money & business
ChatGPT gains voice and vision
A new text-to-speech model, built with professional voice actors, let users talk to ChatGPT and show it photos, ahead of the fully real-time GPT-4o release.
Models & capabilities
xAI unveils Grok
A chatbot with real-time access to X posts and a deliberately irreverent persona, initially released to a small group ahead of a wider subscriber rollout.
Models & capabilities
Anthropic releases Claude 2.1
Claude 2.1 ships with a 200K-token context window, reduced hallucination rates and tool use support.
Models & capabilities
Google launches Gemini
Google said Gemini Ultra beat human experts on the MMLU benchmark; days later Bloomberg reported the model's showcase video had been edited and was not real-time.
Models & capabilities · Culture & impact
Google renames Bard to Gemini and launches Gemini Advanced with Ultra 1.0
The $19.99-a-month Google One AI Premium tier gave access to Ultra 1.0, which Google said was the first model to outperform human experts on MMLU.
Models & capabilities
Gemini 1.5 Pro ships a million-token context window
A mixture-of-experts model that matched Gemini 1.0 Ultra on many tasks at lower compute, offered in limited preview with up to a million tokens of context.
Models & capabilities
Anthropic's Claude 3 takes the frontier from GPT-4
The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
Models & capabilities · Labs & people
Meta releases Llama 3 and puts its assistant everywhere
8B and 70B open-weight models shipped alongside a much larger, still-training 400B+ version, as Meta AI rolled out across Facebook, Instagram, WhatsApp and Messenger.
Open weights & ecosystem · Models & capabilities
OpenAI launches GPT-4o with real-time voice
A single model handling text, vision and audio end to end, with conversational latency — and a voice that led to a public dispute with Scarlett Johansson.
Models & capabilities · Culture & impact
Google puts AI Overviews on search
A Gemini-generated summary above the links, rolled out to hundreds of millions of US users that week with a target of a billion by year end.
Models & capabilities · Culture & impact
Google unveils Project Astra, a universal AI assistant prototype
A prototype, not a product: Google showed a phone-camera assistant with conversational-speed responses but gave no release date beyond 'later this year'.
Models & capabilities
Apple announces Apple Intelligence with OpenAI inside
ChatGPT access was free and optional, required no account, and Apple said queries would not be logged — with other AI providers to be added later.
Models & capabilities · Money & business
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
OpenAI releases a SearchGPT prototype
The prototype answered queries with cited sources drawn from real-time web results and was opened to roughly 10,000 waitlisted testers and select publishers.
Models & capabilities
xAI releases Grok-2
The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
Models & capabilities
Mistral AI relaunches Le Chat with new features
The update added image generation via Black Forest Labs' Flux Ultra, web search, a canvas editor and a paid tier priced at $14.99 a month.
Models & capabilities
xAI releases Grok-3
xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
Models & capabilities · Benchmarks & progress
Google launches AI Mode in Search
The experimental tab used a 'query fan-out' technique to run multiple related searches at once, launching first to opted-in US Google One AI Premium subscribers.
Models & capabilities
Gemini 2.5 Pro takes the lead on reasoning benchmarks
Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
Models & capabilities · Benchmarks & progress
Google I/O puts Gemini into search and ships Veo 3
AI Mode rolled out to all US Search users, and Veo 3 became the first widely-used video model to generate synchronised dialogue and sound effects alongside the picture.
Models & capabilities
Google launches AI Ultra subscription plan at Google I/O
At $249.99 a month — twelve times the existing Google AI Pro tier — the plan bundled early Veo 3 access, 30TB of storage and YouTube Premium.
Money & business · Models & capabilities
Anthropic launches Claude Opus 4 and Claude Sonnet 4
Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
Models & capabilities · Safety & alignment
Anthropic ships Claude Opus 4.1
Anthropic reported 74.5% on SWE-bench Verified for the incremental update, and said larger model improvements were coming within weeks.
Models & capabilities
GPT-5 launches to a backlash over the model it replaced
OpenAI withdrew GPT-4o and other older models the same day; a routing fault made GPT-5 seem weaker, and paying users' objections forced 4o's return.
Models & capabilities · Culture & impact · Benchmarks & progress
Anthropic launches Claude for Chrome browser agent
Anthropic reported unmitigated browser use failed against 23.6% of prompt-injection attacks in testing, falling to 11.2% with its safety measures in place.
Models & capabilities
OpenAI ships the Atlas browser
Built on Chromium and launched first for macOS only, with a paid 'agent mode' able to complete multi-step tasks like bookings and comparisons.
Models & capabilities
OpenAI releases GPT-5.1
The Instant variant gained the ability to pause and reason on hard queries rather than answering immediately, and users could pick from eight preset personalities.
Models & capabilities
Google ships Gemini 3
Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
Models & capabilities · Benchmarks & progress
Sam Altman declares internal 'Code Red' at OpenAI over Gemini 3 competition
Altman told staff to prioritise ChatGPT quality and delay planned advertising and shopping features, weeks after Google's Gemini 3 outperformed OpenAI's models on several benchmarks.
Labs & people
Google launches revamped Gemini Deep Research on Gemini 3 Pro with Interactions API
A new Interactions API lets outside developers embed Google's research agent in their own apps, the first time the tool has been offered outside Google's own products.
Models & capabilities
OpenAI releases GPT-5.2
Released three weeks after Google's Gemini 3 and following a reported internal OpenAI 'code red,' with a claimed 70.9% win rate against professionals on the GDPval benchmark, up from 38.8% for GPT-5.1.
Models & capabilities
OpenAI updates ChatGPT image generation to GPT Image 1.5
OpenAI moved the launch up from a planned January date, part of a competitive scramble after Google's Gemini 3 and its 'Nano Banana Pro' image tool led leaderboards.
Models & capabilities
Google makes Gemini 3 Flash the default model across its products
Priced at $0.50/$3.00 per million tokens, Google reported it ran three times faster than Gemini 2.5 Pro while scoring 33.7% on Humanity's Last Exam, against 37.5% for Gemini 3 Pro.
Models & capabilities
OpenAI ships GPT-5.2-Codex
OpenAI reported an 'unmatched' 56.4% on the SWE-Bench Pro benchmark and 64% on Terminal-Bench 2.0, alongside new defensive-cybersecurity capabilities.
Models & capabilities
Moonshot AI releases Kimi K2.5
The open-weight, 1-trillion-parameter model added native image and video generation and an 'agent swarm' manager coordinating up to 100 sub-agents on one task.
Open weights & ecosystem · Models & capabilities
Anthropic releases Claude Sonnet 4.6
Early testers preferred it to Sonnet 4.5 on coding tasks about 70% of the time, and to the larger Opus 4.5 about 59% of the time, at unchanged Sonnet pricing.
Models & capabilities
Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
OpenAI releases GPT-5.5
Pitched as OpenAI's most agentic model yet, it shipped alongside a $50,000 bug-bounty for jailbreaks that could extract biological-weapons help.
Models & capabilities
Anthropic releases Claude Opus 4.8
The upgrade arrived just 41 days after Opus 4.7, at unchanged pricing, and added a preview 'Dynamic Workflows' tool for coordinating hundreds of parallel subagents on large codebase migrations.
Models & capabilities
SpaceXAI releases Grok 4.5
Built on a 1.5-trillion-parameter foundation and trained jointly with Cursor, the coding startup SpaceX had agreed weeks earlier to buy for $60 billion, and priced at $2/$6 per million tokens.
Models & capabilities
Anthropic launches Claude Opus 5
Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.
Models & capabilities