Benchmarks · Aggregate indices & arenas
LiveBench
How a model performs on recently created questions, graded by objective ground-truth answers rather than human or LLM judgment, so that scores cannot reflect memorised test data and cannot be inflated by a biased judge model.
Abacus.AI, NYU, University of Maryland and collaborators (Colin White, Samuel Dooley, Tom Goldstein, Yann LeCun and others)Released 24 June 2024Live
LiveBench was built to answer a specific worry about every other benchmark on this list: that a widely used test eventually leaks into a model’s training data, so a high score reflects memorisation rather than ability. A team spanning Abacus.AI, NYU and the University of Maryland — including Tom Goldstein and Yann LeCun among its authors — published it in June 2024 with a fix: draw questions from sources too recent to have been trained on, such as that month’s arXiv papers, news articles and maths competitions, refresh the set on a rolling schedule, and grade every answer against an objective ground truth rather than a human or LLM judge’s opinion.
That design also sidesteps a second problem, the one that dogs arenas and LLM-judged benchmarks alike: a judge model can be gamed by style rather than substance. LiveBench’s 23 tasks, spread across reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, are checked against fixed answers instead. The tradeoff is that the whole question set is periodically replaced — the site had moved from its original 2024-06-24 release to a 2026-06-25 edition by the time of writing — which keeps scores comparison-resistant across versions in the way a static benchmark’s are not.
Unlike single-lab benchmarks, LiveBench is genuinely competitive across the field: every major frontier lab appears on its board, and the overall lead has passed back and forth between Anthropic (Claude 3.5 Sonnet at launch), Google (Gemini 2.5 Pro), OpenAI (o3, later GPT-5.5) and back to Anthropic — Claude Fable 5 leads the current 2026-06-25 set at 83.0, narrowly ahead of OpenAI’s GPT-5.6 Sol at 81.0, with strong open-weight entries from Moonshot, Alibaba and Abacus just behind. One subtlety is essential to reading it, though: because each release is a frozen, dated question set and the site re-runs old sets against new models, two “LiveBench” scores are only comparable when they share a release — a model’s number can rise or fall simply because the exam changed. Because each refresh changes the underlying questions, LiveBench functions less as a fixed yardstick than as a standing methodology for building one — a response to benchmark contamination rather than a single benchmark in the usual sense.
The set
23 objective tasks across seven categories — reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following — drawn from sources such as recent math competitions, arXiv papers and news articles. Every task is scored against a ground-truth answer rather than by human or LLM preference. The full question set is refreshed roughly every six months (the latest as of August 2026 dated 2026-06-25) specifically to stay ahead of training-data contamination, and older releases remain browsable for comparison.
Where it stands
Actively maintained with periodic full refreshes rather than one static set, and genuinely multi-lab — every major frontier lab (OpenAI, Anthropic, Google, xAI, DeepSeek, Alibaba, Moonshot, Meta) appears on the board, and the overall lead has changed hands between Anthropic, Google and OpenAI. Because the question set is versioned, scores from different releases are NOT comparable — always pin a score to its release date.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- June 2024Claude 3.5 Sonnet61.2 overall (v1, 2024-06-24 set)The first LiveBench overall leader, on the launch question set.
- March 2025Gemini 2.5 Pro82.4 overall (2024-11-25 set)Led the 2024-11-25 question set. Note: LiveBench re-runs older sets against new models, so a set's board keeps accruing later releases — the number is 'best on that question set to date', not a snapshot of that month.
- April 2025o3 (high)80.7 overall (2025-04-25 set)OpenAI took the lead on the 2025-04-25 question set.
- November 2025Claude Opus 4.580.1 overall (2025-05-30 set)Anthropic back on top on the 2025-05-30 set.
- April 2026GPT-5.5 (xhigh)80.7 overall (2026-01-08 set)
- June 2026Claude Fable 5 (Max Effort)83.0 overall (2026-06-25 set)Current leader; GPT-5.6 Sol 81.0, with strong open-weight entries (Moonshot Kimi K3, Alibaba Qwen, Abacus Smaug) close behind.
Current best: Claude Fable 5 (Max Effort) — 83.0 overall Overall score on the 2026-06-25 question-set release; GPT-5.6 Sol (Max Effort) followed at 81.0 and GPT-5.5 Thinking (xHigh) at 80.2. Overall = the mean of the seven category averages. Scores are only comparable within the same question-set release.
In the timeline · 2 entries
Alibaba releases Qwen2.5-Max
Unlike most of Alibaba's Qwen line, Max was released as a proprietary API-only model, pretrained on over 20 trillion tokens, which Alibaba said beat DeepSeek-V3 on several benchmarks.
Models & capabilities · Benchmarks & progress
StepFun launches Step-2, a trillion-parameter MoE model
StepFun said the mixture-of-experts model approximated GPT-4 on maths, logic, coding and dialogue; it was unveiled alongside a multimodal and an image-generation model at WAIC.
Models & capabilities · Open weights & ecosystem