Benchmarks · Coding & software engineering

LiveCodeBench

also: LCB

How well a model codes on problems it could not have memorised, by dating every problem and checking performance separately on those published before and after the model's training cutoff.

UC Berkeley, MIT & CornellReleased 12 March 2024Live

By early 2024, two of coding’s most-cited benchmarks had a shared problem: HumanEval and MBPP had been sitting on the open web for years, and nobody could rule out that their exact problems and solutions had been absorbed into the training data of the models being scored against them. LiveCodeBench, published by researchers at UC Berkeley, MIT and Cornell, answered with a benchmark that could not go stale in the same way: it continuously pulls fresh problems from LeetCode, AtCoder and Codeforces contests and tags each with its publication date, so a model’s score can be split between problems that predate its training cutoff and ones that came after.

The date-segmented design did what it was built to do — it produced direct evidence that some models scored measurably better on older, plausibly-memorised problems than on newer ones, turning contamination from a suspicion into something a benchmark could actually show. Beyond plain code generation, LiveCodeBench also grades self-repair, test-output prediction and code execution, giving a fuller picture of coding competence than a single pass/fail generation task. It was adopted quickly and is now commonly reported across major labs’ model cards — DeepSeek, Google and Alibaba’s Qwen all cite it, though Anthropic and OpenAI tend to lead their own headline tables with SWE-bench instead.

That refreshing pool is also what makes LiveCodeBench numbers easy to misread: because the benchmark fixes a May-2023 start and only extends the end date with each version (v6 reaches April 2025), two “LiveCodeBench” scores are comparable only if they cover the same window — the same Gemini 2.5 Pro build scored 75.6% on one window and 69.0% on a later one. On the benchmark’s own official cross-lab leaderboard, for the August-2024–May-2025 window, OpenAI’s o4-mini led at 80.2%, ahead of DeepSeek’s R1-0528 at 73.1% and Gemini 2.5 Pro at 73.6% — a reminder that it is a genuinely multi-lab board, not a DeepSeek showcase. That official board has not since added the late-2026 frontier; independent third-party runs put the current top models (Claude Fable 5, Claude Opus 5, Gemini 3.1 Pro) near 90% on v6.

Because the problem pool keeps refreshing, LiveCodeBench has so far avoided the saturation that overtook its predecessors, and its live-update approach became a template other domains borrowed when they ran into the same contamination problem. It is not the final word on coding ability — the underlying problems are still competitive-programming puzzles rather than the messy, multi-file work of real software engineering, which is closer to what benchmarks like SWE-bench try to test — but as a memorisation-resistant coding number, it has become a common citation.

The set

Competitive-programming problems continuously scraped from LeetCode, AtCoder and Codeforces contests, each tagged by publication date. Beyond plain code generation, it also scores self-repair (fixing code from a failing test), test-output prediction, and code execution — a broader slice of coding ability than a single generate-and-check pass.

Example

One problem in the set, titled 'Anti', verbatim: 'A DDoS-type string is a string of length 4... The first and second characters are equal. For instance, DDoS and AAaA are DDoS-type strings.' Given a string with wildcards, count completions avoiding a DDoS-type subsequence, mod 998244353.huggingface.co

Where it stands

A contamination-resistant coding benchmark widely reported across major labs' model cards (DeepSeek, Google, Alibaba/Qwen), though Anthropic and OpenAI tend to lead their own tables with SWE-bench. Scores are only comparable within the same problem window — the benchmark fixes a May-2023 start and extends the end date each version (v6 runs to April 2025). Its official leaderboard is currently frozen around mid-2025 models (o4-mini tops it at 80.2% on the Aug-2024–May-2025 window); current-frontier numbers come from independent third-party runs, where late-2026 models cluster near 90%.

How the top score changed hands

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. March 2025DeepSeek V3-032449.2%Up from 39.2% for the preceding DeepSeek V3 checkpoint, per DeepSeek's own release notes — an early, non-reasoning self-report.
  2. April 2025o4-mini (High)80.2%Top of LiveCodeBench's own official cross-lab leaderboard on the fixed Aug-2024–May-2025 window (2408-2505); an independently-computed board on which DeepSeek-R1-0528 scored 73.1% and Gemini 2.5 Pro 73.6% on the same window.
  3. June 2026Claude Fable 589.8% (vals.ai v6 run)The current frontier, via the independent vals.ai LiveCodeBench v6 run (Aug 2026); Opus 5, Gemini 3.1 Pro and GPT-5.2 Codex all within ~2 points.

Current best: Claude Fable 5 — 89.8% (vals.ai v6 run) Code-generation Pass@1 on the independent third-party vals.ai LiveCodeBench v6 run (Aug 2026); Claude Opus 5 (89.0%), Gemini 3.1 Pro (88.5%) and GPT-5.2 Codex (88.0%) cluster just behind. LiveCodeBench's own official leaderboard has not added late-2026 frontier models. Note that v6 problems end ~April 2025, before these models' training cutoffs, so contamination-resistance is weaker for them than the benchmark's design intends.

In the timeline · 9 entries

More coding & software engineering benchmarks