Benchmarks · Coding & software engineering
LiveCodeBench
also: LCB
How well a model codes on problems it could not have memorised, by dating every problem and checking performance separately on those published before and after the model's training cutoff.
UC Berkeley, MIT & CornellReleased 12 March 2024Live
By early 2024, two of coding’s most-cited benchmarks had a shared problem: HumanEval and MBPP had been sitting on the open web for years, and nobody could rule out that their exact problems and solutions had been absorbed into the training data of the models being scored against them. LiveCodeBench, published by researchers at UC Berkeley, MIT and Cornell, answered with a benchmark that could not go stale in the same way: it continuously pulls fresh problems from LeetCode, AtCoder and Codeforces contests and tags each with its publication date, so a model’s score can be split between problems that predate its training cutoff and ones that came after.
The date-segmented design did what it was built to do — it produced direct evidence that some models scored measurably better on older, plausibly-memorised problems than on newer ones, turning contamination from a suspicion into something a benchmark could actually show. Beyond plain code generation, LiveCodeBench also grades self-repair, test-output prediction and code execution, giving a fuller picture of coding competence than a single pass/fail generation task. It was adopted quickly and is now commonly reported across major labs’ model cards — DeepSeek, Google and Alibaba’s Qwen all cite it, though Anthropic and OpenAI tend to lead their own headline tables with SWE-bench instead.
That refreshing pool is also what makes LiveCodeBench numbers easy to misread: because the benchmark fixes a May-2023 start and only extends the end date with each version (v6 reaches April 2025), two “LiveCodeBench” scores are comparable only if they cover the same window — the same Gemini 2.5 Pro build scored 75.6% on one window and 69.0% on a later one. On the benchmark’s own official cross-lab leaderboard, for the August-2024–May-2025 window, OpenAI’s o4-mini led at 80.2%, ahead of DeepSeek’s R1-0528 at 73.1% and Gemini 2.5 Pro at 73.6% — a reminder that it is a genuinely multi-lab board, not a DeepSeek showcase. That official board has not since added the late-2026 frontier; independent third-party runs put the current top models (Claude Fable 5, Claude Opus 5, Gemini 3.1 Pro) near 90% on v6.
Because the problem pool keeps refreshing, LiveCodeBench has so far avoided the saturation that overtook its predecessors, and its live-update approach became a template other domains borrowed when they ran into the same contamination problem. It is not the final word on coding ability — the underlying problems are still competitive-programming puzzles rather than the messy, multi-file work of real software engineering, which is closer to what benchmarks like SWE-bench try to test — but as a memorisation-resistant coding number, it has become a common citation.
The set
Competitive-programming problems continuously scraped from LeetCode, AtCoder and Codeforces contests, each tagged by publication date. Beyond plain code generation, it also scores self-repair (fixing code from a failing test), test-output prediction, and code execution — a broader slice of coding ability than a single generate-and-check pass.
Example
One problem in the set, titled 'Anti', verbatim: 'A DDoS-type string is a string of length 4... The first and second characters are equal. For instance, DDoS and AAaA are DDoS-type strings.' Given a string with wildcards, count completions avoiding a DDoS-type subsequence, mod 998244353.huggingface.co
Where it stands
A contamination-resistant coding benchmark widely reported across major labs' model cards (DeepSeek, Google, Alibaba/Qwen), though Anthropic and OpenAI tend to lead their own tables with SWE-bench. Scores are only comparable within the same problem window — the benchmark fixes a May-2023 start and extends the end date each version (v6 runs to April 2025). Its official leaderboard is currently frozen around mid-2025 models (o4-mini tops it at 80.2% on the Aug-2024–May-2025 window); current-frontier numbers come from independent third-party runs, where late-2026 models cluster near 90%.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- March 2025DeepSeek V3-032449.2%Up from 39.2% for the preceding DeepSeek V3 checkpoint, per DeepSeek's own release notes — an early, non-reasoning self-report.
- April 2025o4-mini (High)80.2%Top of LiveCodeBench's own official cross-lab leaderboard on the fixed Aug-2024–May-2025 window (2408-2505); an independently-computed board on which DeepSeek-R1-0528 scored 73.1% and Gemini 2.5 Pro 73.6% on the same window.
- June 2026Claude Fable 589.8% (vals.ai v6 run)The current frontier, via the independent vals.ai LiveCodeBench v6 run (Aug 2026); Opus 5, Gemini 3.1 Pro and GPT-5.2 Codex all within ~2 points.
Current best: Claude Fable 5 — 89.8% (vals.ai v6 run) Code-generation Pass@1 on the independent third-party vals.ai LiveCodeBench v6 run (Aug 2026); Claude Opus 5 (89.0%), Gemini 3.1 Pro (88.5%) and GPT-5.2 Codex (88.0%) cluster just behind. LiveCodeBench's own official leaderboard has not added late-2026 frontier models. Note that v6 problems end ~April 2025, before these models' training cutoffs, so contamination-resistance is weaker for them than the benchmark's design intends.
In the timeline · 9 entries
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
Moonshot AI releases Kimi K2, a 1-trillion-parameter open-weight model
The mixture-of-experts model activates 32 billion of its 1 trillion parameters per token and was trained with the Muon optimiser at a scale its makers said had previously caused instability.
Open weights & ecosystem · Models & capabilities
DeepSeek releases DeepSeek-R1-0528 update
Released under an MIT licence, the update raised AIME 2025 accuracy from 70% to 87.5% by roughly doubling the average length of the model's reasoning traces.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
DeepSeek releases DeepSeek-V3-0324 update
The updated checkpoint scored 81.2% on MMLU-Pro and 59.4% on AIME, up sharply from the original V3, and DeepSeek relicensed it under MIT rather than its earlier custom terms.
Open weights & ecosystem · Models & capabilities
Alibaba releases Qwen2.5-Max
Unlike most of Alibaba's Qwen line, Max was released as a proprietary API-only model, pretrained on over 20 trillion tokens, which Alibaba said beat DeepSeek-V3 on several benchmarks.
Models & capabilities · Benchmarks & progress
Moonshot AI releases Kimi K1.5 reasoning model
Moonshot said its RL-trained model matched OpenAI's o1 on multimodal reasoning without Monte Carlo tree search, but it launched the same week as DeepSeek-R1 and drew far less attention.
Models & capabilities · Benchmarks & progress
OpenAI ships o1 model with new developer tools
The full o1 reasoning model reached the API alongside function calling, structured outputs and vision support for developers.
Models & capabilities
Alibaba releases Qwen2.5-Coder
Open-weight coding-specialised model family built on Qwen2.5, aimed at competing with DeepSeek-Coder and closed coding models.
Open weights & ecosystem · Models & capabilities
LiveCodeBench paper published
Testing 52 models against problems tagged by publication date, the authors found evidence that some scored higher on problems predating their training cutoff.
Benchmarks & progress