Benchmarks · Aggregate indices & arenas
LMArena
also: Chatbot Arena, LMSYS Chatbot Arena, Arena
Which of two anonymous models people prefer in open-ended, head-to-head conversation, aggregated into a running Elo rating rather than a fixed test score.
LMSYS / UC Berkeley, later independent as LMArena (Arena Intelligence)Released 3 May 2023Disputed
Most benchmarks score a model against a fixed answer key. LMArena, launched by LMSYS researchers at UC Berkeley as Chatbot Arena in May 2023, asks a different question: given two anonymous replies to the same prompt, which one does a person actually prefer? Visitors vote blind, the identities are revealed afterwards, and the votes accumulate into a running Elo-style rating rather than a one-off score — closer to how the models are used in practice, but harder to audit than a test with a right answer.
The approach made LMArena one of the most closely watched public rankings in the field, cited in launch announcements from every major lab as the votes piled into the millions. That prominence also made it a target. In April 2025 it emerged that Meta had submitted a specially tuned Llama 4 Maverick variant that outperformed the model it actually shipped, and weeks later “The Leaderboard Illusion” argued that undisclosed private pre-release testing structurally favoured well-resourced labs. LMArena disputed several of the paper’s figures but adopted disclosure and provisional-scoring policies in response.
Leadership has continued to change hands quickly since — Google’s Gemini 2.5 Pro and later Gemini 3 Pro both took the top spot on release, and by August 2026 the text arena, now rebranded simply “Arena,” had logged more than 7.7 million votes across 389 models, with Anthropic, Google and OpenAI systems clustered within a few Elo points at the top. It remains a widely cited signal of how models are received in open-ended use, but after 2025’s disputes it is generally read alongside fixed benchmarks rather than in place of them.
The set
Visitors submit a prompt to two randomly paired, anonymised models, read both replies and vote for the better one before either model's identity is revealed. Votes accumulate into a Bradley-Terry/Elo-style rating; as of August 2026 the text arena alone has logged over 7.7 million votes across 389 models.
Where it stands
Still one of the most-cited public rankings, but 2025's 'Leaderboard Illusion' paper and the Llama 4 Maverick episode left its scores widely treated as suggestive rather than authoritative.
How the top score changed hands
- May 2023Vicuna-13Bleader of the initial ~4,700-vote rankingThe launch leaderboard covered nine mostly open-weight, LLaMA-derived models; frontier proprietary models were not yet included.
- March 2025Gemini 2.5 Proled by what Google called 'a significant margin'Google's own release claim; came after roughly two years in which OpenAI and Anthropic models had generally led the board.
- November 2025Gemini 3 Pro1501 EloReported alongside record scores on GPQA Diamond and SWE-bench Verified at launch.
Current best: Claude Fable 5 — 1506 Elo Overall text leaderboard snapshot; rankings shift week to week and several models cluster within a few Elo points of the top.
In the timeline · 22 entries · showing 16 most notable
GLM-5.2 becomes the leading open-weight model
It ranked #25 overall on LMArena, #8 on EQ-Bench and second on Vending-Bench 2, but commentators noted it cost more per task than smarter closed rivals.
Open weights & ecosystem · Models & capabilities
Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal model
Baidu said the mixture-of-experts model activates under 3% of its parameters per query and ranked first among Chinese models, eighth globally, on LMArena's text leaderboard.
Models & capabilities · Benchmarks & progress
Mistral launches Mistral 3 model family
Mistral released Mistral 3, including dense models at 3B/8B/14B and a new mixture-of-experts Mistral Large 3 (41B active, 675B total), all under Apache 2.0.
Models & capabilities · Open weights & ecosystem
Google ships Gemini 3
Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
Models & capabilities · Benchmarks & progress
xAI releases Grok 4.1
xAI tuned the update for personality and reliability rather than raw reasoning, reporting a two-week blind test in which users preferred it to Grok 4 64.8% of the time.
Models & capabilities
Design Arena launches as crowdsourced AI design benchmark
The Y Combinator-backed site shows visitors two AI-generated designs from an identical prompt and asks them to pick the better one, ranking models by Elo-style score.
Benchmarks & progress
Google launches Kaggle Game Arena AI chess tournament
Eight frontier models played an all-play-all chess tournament of over 100 matches, with the game harness open-sourced so the contest could be independently verified.
Benchmarks & progress
Google updates Gemini 2.5 Pro preview with improved coding performance
The update, internally labelled 06-05, also led coding benchmarks including Aider Polyglot and performed strongly on Humanity's Last Exam.
Models & capabilities · Benchmarks & progress
LMArena responds to 'Leaderboard Illusion' paper with policy changes
LMArena disputed the paper's headline figures on open-model share and score-boosting but agreed to mark scores 'provisional' and disclose pre-release testing.
Benchmarks & progress
'The Leaderboard Illusion' paper critiques Chatbot Arena methodology
Researchers found Meta tested roughly 27 private Llama variants before its public release and that OpenAI and Google alone received about 40% of all Arena battle data.
Benchmarks & progress
Meta accused of gaming LMArena with tuned Llama 4 Maverick variant
The version ranked second on the leaderboard, labelled 'Llama-4-Maverick-03-26-Experimental', produced longer, emoji-heavy answers than the model Meta actually shipped for download.
Benchmarks & progress · Open weights & ecosystem
Llama 4 lands badly
Meta's mixture-of-experts release was undercut by accusations that a version tuned for LMArena differed from the public weights.
Open weights & ecosystem · Benchmarks & progress · Models & capabilities
Gemini 2.5 Pro takes the lead on reasoning benchmarks
Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
Models & capabilities · Benchmarks & progress
Google releases Gemma 3, an open model family built on Gemini 2.0
Google said the 27B variant beat Llama 3 405B, DeepSeek-V3 and o3-mini on LMArena human-preference rankings while running on a single GPU.
Open weights & ecosystem · Models & capabilities
xAI releases Grok-3
xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
Models & capabilities · Benchmarks & progress
LMSYS launches Chatbot Arena
The Berkeley-linked project ranked chatbots by anonymous, randomised head-to-head votes rather than fixed test sets, and later became LMArena.
Benchmarks & progress