Reference

AI benchmarks

A benchmark is a fixed test that lets different AI systems be compared on the same footing. This catalogue is ordered by how live each one is: the benchmarks the frontier labs actively compete on — many scores, from many labs, with the lead still changing hands — sit largest and first, marked 🔥 Competitive; the ones that were an interesting idea but never grew a live leaderboard trail off smaller below. The recurring cycle — a benchmark is built to be hard, models climb it within a year or two, it saturates, a harder one replaces it — is followed as a story in the benchmarks and their saturation thread.

118 benchmarks · 12 domains · 32 competitive

Coding & software engineering

Writing code, fixing real bugs, resolving issues in live repositories.

  • DeepSWE🔥 CompetitiveLive

    Can a coding agent complete an original, long-horizon software engineering task in a real repository, graded by whether the behaviour is correct — not just whether it matches one specific reference implementation?

    Datacurve8 July 2026Leader: Gemini 3.6 Flash (49% (DeepSWE v1.1))

  • SWE-bench Pro🔥 CompetitiveLive

    Can a model resolve a realistic, multi-file software engineering task in a codebase it could not have memorised — including private, commercial code rather than only well-known open-source repositories?

    Scale AI19 September 2025Leader: Muse Spark 1.1 (Meta) (61.5% ± 3.1)

  • SWE-bench🔥 CompetitiveSaturating

    Can a model resolve a real, unseen GitHub issue by editing a codebase so that the project's own hidden tests pass?

    Princeton & Stanford10 October 2023Leader: Claude Opus 4.8 (88.6%)

  • Codeforces / CodeContests🔥 CompetitiveLive

    Can a model solve genuinely novel algorithmic problems under contest conditions — the kind that require devising an approach, not recalling one — well enough to rank against real competitive programmers?

    Codeforces (competitive-programming platform); CodeContests dataset curated by Google DeepMind2 February 2022Leader: Gemini 3 Deep Think (v2) (Elo 3455)

  • Terminal-Bench🔥 CompetitiveLive

    Can an AI agent actually operate a computer through a real command-line shell — issuing commands, reading their output, and adapting — to finish a multi-step task, rather than just producing plausible-looking commands?

    Stanford University & Laude Institute19 May 2025

  • LiveCodeBenchLive

    How well a model codes on problems it could not have memorised, by dating every problem and checking performance separately on those published before and after the model's training cutoff.

    UC Berkeley, MIT & Cornell12 March 2024Leader: Claude Fable 5 (89.8% (vals.ai v6 run))

  • HumanEvalSaturated

    Can a model write a correct, working Python function from a natural-language docstring alone?

    OpenAI7 July 2021Leader: Claude 3.5 Sonnet (92.0%)

  • Aider PolyglotLive

    Can a model act as a practical pair-programmer — reading an existing multi-language codebase, understanding a task, and editing the actual files correctly, not just writing an isolated function?

    Paul Gauthier / Aider21 December 2024Leader: GPT-5 (high reasoning effort) (88.0%)

  • SWE-LancerLive

    OpenAI18 February 2025Leader: GPT-5.1-Codex-Max (79.9%)

  • NanoGPT Speedrun🔥 CompetitiveLive

    Keller Jordan & open community; formalised into an agent benchmark by Meta FAIR researchers28 May 2024

  • MBPPSaturated

    Google Research (Austin et al.)16 August 2021

  • BigCodeBenchLive

    BigCode Project (Hugging Face, ServiceNow & collaborators)22 June 2024

  • Konwinski PrizeRetired

    Andy Konwinski12 December 2024

Reasoning & problem-solving

General reasoning, hard exams, puzzles that resist memorisation.

  • GPQA🔥 CompetitiveSaturating

    Whether a model can answer graduate-level science questions that a skilled non-expert cannot solve even with unrestricted web access and half an hour per question.

    NYU, Cohere & Anthropic researchers20 November 2023Leader: Gemini 3 Pro (Deep Think) (93.8%)

  • Humanity's Last Exam🔥 CompetitiveLive

    Whether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.

    Center for AI Safety & Scale AI23 January 2025

  • EnigmaEval🔥 CompetitiveLive

    Long-horizon multimodal reasoning: synthesising implicit knowledge and chaining many steps of lateral deduction over mixed image-and-text puzzle-hunt problems that take skilled human teams hours to days to solve.

    Scale AI & Center for AI Safety13 February 2025Leader: Claude Opus 5 (43.9%)

  • ARC-AGI🔥 CompetitiveLive

    Whether a system can infer an unfamiliar abstract rule from a handful of examples and apply it to a new case, rather than recognising a pattern it has seen before.

    ARC Prize Foundation (François Chollet, Mike Knoop)5 November 2019

  • MMLU-Pro🔥 CompetitiveLive

    Whether a model's knowledge holds up once guessing is made hard and questions demand multi-step reasoning rather than recall.

    TIGER-AI-Lab (Wang, Ma, Zhang et al.)3 June 2024Leader: Gemini 3.1 Pro (91.16%)

  • SimpleBenchLive

    SimpleBench TeamOctober 2024Leader: Claude Fable (81.9%)

  • DROPRetired

    AI2 & UC Irvine (Dua, Wang, Dasigi, Gardner et al.)1 March 2019Leader: GPT-4 (80.9 F1)

  • BIG-BenchRetired

    450+ contributors across 132 institutions (Google-led collaboration)9 June 2022

  • AGIEvalSaturated

    Microsoft Research (Zhong, Duan, Chen et al.)13 April 2023Leader: GPT-4o (71.4% (English tasks, few-shot))

  • ARC (AI2 Reasoning Challenge)Retired

    Allen Institute for AI (AI2)14 March 2018

  • HellaSwagSaturated

    University of Washington & Allen Institute for AI (Zellers et al.)19 May 2019Leader: GPT-4 (base, 10-shot) (95.3%)

  • WinoGrandeRetired

    Allen Institute for AI (Sakaguchi, Le Bras, Bhagavatula, Choi)24 July 2019

Multimodal

Vision, audio, video and charts — reasoning over more than text.

  • MMMU🔥 CompetitiveLive

    Can a model answer college-exam-level questions that genuinely require reading an accompanying image — a chart, diagram, map or chemical structure — rather than knowledge alone?

    Ohio State University & University of Waterloo (Yue, Su, Chen et al.)27 November 2023Leader: Gemini 3 Flash (81.2% (MMMU-Pro))

  • MindCube🔥 CompetitiveLive

    Whether a vision-language model can build a spatial mental model of a scene — inferring the positions, orientations and possible movements of objects, including ones it cannot currently see — from only a few limited views.

    Northwestern University (MLL-Lab), with Stanford, NYU & University of Washington26 June 2025Leader: Claude Opus 5 (84.3%)

  • SpatialViz🔥 CompetitiveLive

    Spatial visualization: whether a multimodal model can mentally imagine and manipulate visual structures that are not directly observable, across mental rotation, mental folding, visual penetration and mental animation.

    Institute of Automation, Chinese Academy of Sciences (CASIA), with UCAS, SJTU, ShanghaiTech, Huawei Noah's Ark & UCL10 July 2025Leader: GPT-5.6 Sol (80.7%)

  • ERQA🔥 CompetitiveLive

    Embodied Reasoning QA: whether a vision-language model understands a physical scene well enough to reason about acting in it — spatial relations, trajectories, state estimation, pointing and multi-view correspondence.

    Google DeepMind (Gemini Robotics Team)12 March 2025Leader: Gemini 3.5 Flash (75.4%)

  • IntPhys 2🔥 CompetitiveLive

    Intuitive physics from video: whether a model grasps four macroscopic principles — object permanence, immutability, spatio-temporal continuity and solidity — well enough to tell physically possible events from impossible ones.

    FAIR at Meta11 June 2025Leader: Claude Fable 5 (67.1%)

  • MathVistaSaturating

    UCLA, University of Washington & Microsoft Research (Lu, Bansal, Galley, Gao et al.)3 October 2023Leader: o3 (86.8% (testmini))

  • Video-MMELive

    Nanjing University, with XMU, HKU, PKU, CUHK, ECNU & CASIA31 May 2024Leader: Gemini 1.5 Pro (81.3% (with subtitles))

  • ChartQASaturating

    York University with NTU Singapore & Salesforce Research (Masry, Hoque, Joty et al.)19 March 2022Leader: Claude 3.5 Sonnet (90.8% (test, relaxed accuracy))

  • BLINKLive

    University of Pennsylvania, University of Washington, Allen Institute for AI, UC Davis and Columbia18 April 2024Leader: GPT-4o (68.0%)

  • RealWorldQALive

    xAIApril 2024Leader: InternVL2.5-78B (78.7%)

  • MVBenchLive

    OpenGVLab, Shanghai AI Laboratory, with Nanjing University, Fudan and University of Hong Kong28 November 2023Leader: Qwen2.5-VL-72B (70.4%)

  • AI2DSaturating

    Allen Institute for AI & University of Washington24 March 2016Leader: InternVL2.5-78B (89.1% (with mask))

  • DocVQASaturated

    IIIT Hyderabad, CVC (Universitat Autònoma de Barcelona) & Amazon1 July 2020Leader: Qwen2.5-VL-72B (96.4% (test, ANLS))

  • MMBenchSaturating

    OpenCompass / Shanghai AI Laboratory (Liu, Duan, Chen, Lin et al.)12 July 2023Leader: Qwen2.5-VL-72B (88.4% (MMBench-V1.1-EN test))

  • ScreenSpotSaturating

    Nanjing University, with Shanghai AI Laboratory and National University of Singapore17 January 2024Leader: SeeClick (53.4%)

  • VQAv2Saturated

    Virginia Tech, Georgia Tech and US Army Research Laboratory (Goyal et al.)2 December 2016

Agents, tools & computer use

Multi-step tasks: driving a browser, a terminal, a desktop, real tools.

  • AutomationBench🔥 CompetitiveLive

    Can an AI agent carry out a real business workflow across multiple SaaS applications — finding the right API endpoints itself, following a company's own layered business rules, and getting the right data to the right system — rather than completing a single, well-specified task?

    Zapier21 April 2026Leader: GPT-6 Astra (41.4%)

  • TextQuests🔥 CompetitiveLive

    Long-horizon agentic reasoning: whether an LLM can autonomously play classic text-adventure games, sustaining exploration, planning and memory over a long and growing context without external tools.

    Center for AI Safety, with Carnegie Mellon University & Gray Swan AI31 July 2025Leader: Claude Fable 5 (56.1%)

  • OSWorld🔥 CompetitiveLive

    Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?

    University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of Waterloo11 April 2024

  • BrowseComp🔥 CompetitiveDisputed

    Can an agent find a specific, hard-to-locate fact on the open web by searching persistently and connecting scattered clues, rather than by knowing the answer already or finding it in one search?

    OpenAI10 April 2025Leader: MiniMax-M3 (83.5%)

  • τ-benchLive

    Sierra17 June 2024Leader: Qwen3.5-397B-A17B (87.9%)

  • GAIALive

    Meta AI (FAIR), Hugging Face & AutoGPT21 November 2023Leader: HAL Generalist Agent (Claude Sonnet 4.5) (74.55%)

  • VisualWebArenaLive

    Carnegie Mellon University24 January 2024Leader: Gemini 2.5 Flash (54.0%)

  • WebVoyagerLive

    Tencent AI Lab & Westlake University25 January 2024Leader: OpenAI Operator (CUA) (87%)

  • TheAgentCompanyLive

    Carnegie Mellon University18 December 2024Leader: Gemini 2.5 Pro (30.3% (full completion) / 39.3% (partial credit))

  • AgentBenchRetired

    Tsinghua University, with Ohio State University & UC Berkeley7 August 2023

  • WebArenaLive

    Carnegie Mellon University25 July 2023Leader: IBM CUGA (61.7%)

  • AndroidWorldSaturating

    Google Research / Google DeepMind23 May 2024Leader: AGI-0 (97.4%)

  • Online-Mind2WebDisputed

    Ohio State University NLP group, with UC Berkeley2 April 2025Leader: OpenAI Operator (61.3%)

Aggregate indices & arenas

Human-preference arenas and composite indices that roll many tests into one ranking.

  • LMArena🔥 CompetitiveDisputed

    Which of two anonymous models people prefer in open-ended, head-to-head conversation, aggregated into a running Elo rating rather than a fixed test score.

    LMSYS / UC Berkeley, later independent as LMArena (Arena Intelligence)3 May 2023Leader: Claude Fable 5 (1506 Elo)

  • Artificial Analysis Intelligence Index🔥 CompetitiveLive

    A single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.

    Artificial AnalysisQ1 2024

  • LiveBenchLive

    How a model performs on recently created questions, graded by objective ground-truth answers rather than human or LLM judgment, so that scores cannot reflect memorised test data and cannot be inflated by a biased judge model.

    Abacus.AI, NYU, University of Maryland and collaborators (Colin White, Samuel Dooley, Tom Goldstein, Yann LeCun and others)24 June 2024Leader: Claude Fable 5 (Max Effort) (83.0 overall)

  • Artificial Analysis Coding Agent Index🔥 CompetitiveLive

    How well a coding agent — a specific model paired with a specific harness, such as Claude Code or Codex — completes real software-engineering work end to end, not just whether the underlying model answers a coding question correctly.

    Artificial AnalysisMay 2026Leader: Claude Code – Claude Opus 5 (xhigh) / Codex – GPT-5.6 Sol (max) (67 (tied))

  • HELMLive

    Stanford CRFM (Center for Research on Foundation Models, part of Stanford HAI)16 November 2022Leader: GPT-5 mini (2025-08-07) (0.819 mean score)

  • Open LLM LeaderboardRetired

    Hugging FaceQ2 2023

Mathematics

From grade-school word problems to unsolved research mathematics.

  • AIME🔥 CompetitiveContamination concerns

    Whether a model can solve short, single-answer competition-maths problems requiring several steps of reasoning but no written proof.

    Mathematical Association of America (the underlying competition); repurposed as an LLM benchmark by the reasoning-model research community from late 2024September 2024Leader: GPT-5.2 (100.0%)

  • MATH🔥 CompetitiveSaturated

    Can a model solve a competition-level mathematics problem and produce a correct step-by-step derivation, not just a lucky final number?

    UC Berkeley (Hendrycks et al.)5 March 2021Leader: Kimi K2 (97.4%)

  • FrontierMath🔥 CompetitiveDisputed

    Can a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?

    Epoch AI8 November 2024Leader: GPT-5.2 (40.3% (Tiers 1-3))

  • GSM8KSaturated

    Can a model solve a grade-school arithmetic word problem that takes several linked steps to work through, rather than a single calculation?

    OpenAI27 October 2021Leader: Llama 3.1 405B (96.8%)

  • MathArenaLive

    ETH Zurich SRI Lab & INSAITMay 2025Leader: Claude Opus 5 (max) (84.4% ± 2.8% (overall expected performance))

  • miniF2FSaturated

    OpenAI (Zheng, Han & Polu)31 August 2021Leader: Goedel-Prover-V2-32B (90.4% (self-correction), 88.1% (standard), pass@32)

  • HARPLive

    Albert S. Yue, Lovish Madaan, Ted Moskovitz, DJ Strouse and Aaditya K. Singh11 December 2024Leader: o1-mini (41.1%)

  • PutnamBenchLive

    UT Austin (Tsoukalas, Chaudhuri et al.)15 July 2024Leader: Seed-Prover 1.5 (ByteDance) (88%)

  • HMMTLive

    Harvard and MIT undergraduates (the underlying competition); evaluated as a live LLM benchmark by ETH Zurich's MathArena projectFebruary 2025Leader: Inkling-Small (90.2%)

  • Omni-MATHLive

    Peking University & Alibaba10 October 2024Leader: OpenAI o1-mini (60.54%)

Real-world & economic value

Professional work products and economically valuable tasks.

  • GDPval🔥 CompetitiveLive

    Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.

    OpenAI25 September 2025

  • Agents' Last Exam🔥 CompetitiveLive

    Whether an AI agent can complete long-horizon, economically valuable professional tasks — not just answer questions — with a verifiable, checkable outcome.

    UC Berkeley RDI (Dawn Song et al.)3 June 2026Leader: GPT-5.6 Sol (OpenAI, Codex harness) (30.6% pass rate)

  • Vending-BenchLive

    Andon Labs20 February 2025Leader: Claude Opus 4.7 (first place on Vending-Bench 2)

  • WorkBenchLive

    Olly Styles et al. (University of Warwick); maintained by MindsDB1 May 2024Leader: Claude Fable 5 (98% task completion, harmful action on 1.9% of tasks)

  • METR Time HorizonLive

    METR19 March 2025Leader: Claude Opus 4.5 (~320 minutes)

  • CORE-BenchLive

    Princeton University17 September 2024Leader: CORE-Agent (GPT-4o) (45.9% overall (21% on the hardest tier))

Safety, security & robustness

Jailbreaks, dangerous-capability evals, honesty, sabotage, harm.

  • ExploitBench🔥 CompetitiveLive

    How far an AI agent gets through the actual chain of an exploit — not just whether it crashes a target, but whether it can turn that crash into control of the machine.

    Carnegie Mellon University13 May 2026Leader: GPT-6 Astra (100.0% (capture rate))

  • ExploitGym🔥 CompetitiveLive

    Can an AI agent turn a known software vulnerability into a real, working attack — not merely identify or patch it, but exploit it end to end, including against active defences?

    Google, UC Berkeley, MPI-SP, UCSB & collaborators11 May 2026Leader: Claude Mythos Preview (157 / 898 instances)

  • SEC-bench Pro🔥 CompetitiveLive

    Can a model find a genuine, previously undisclosed-style vulnerability in a large, real codebase and prove it with a working exploit input — not just patch a bug it has already been shown?

    University of Illinois Urbana-Champaign & UC Berkeley26 May 2026Leader: Codex (GPT-5.5) (58%)

  • MASKLive

    Center for AI Safety & Scale AI5 March 2025Leader: Claude 3.7 Sonnet (73.4% honesty (26.6% lying rate))

  • CyberGymLive

    UC Berkeley (Sunblaze Lab / Berkeley RDI)3 June 2025Leader: GLM-5.3 (84.5% (reproduction rate))

  • AgentHarmLive

    UK AI Security Institute, Gray Swan AI & collaborators11 October 2024

  • SHADE-ArenaLive

    Anthropic16 June 2025

  • HarmBenchLive

    UIUC, Center for AI Safety & collaborators6 February 2024

  • StrongREJECTLive

    UC Berkeley, Center for Human-Compatible AI15 February 2024

  • CybenchLive

    Stanford University15 August 2024

  • CyberSecEvalLive

    Meta AI7 December 2023

  • WMDPLive

    Center for AI Safety & a consortium including UC Berkeley, MIT, Scale AI and SecureBio5 March 2024

  • JailbreakBenchLive

    University of Pennsylvania, EPFL & collaborators28 March 2024

  • SEC-benchLive

    University of Illinois Urbana-Champaign & Purdue13 June 2025

Science & research

Domain knowledge and research work in the sciences.

  • HealthBench🔥 CompetitiveLive

    How well does a model handle realistic, open-ended health conversations — with a layperson or a clinician — judged against criteria that practising physicians say actually matter, rather than a multiple-choice medical exam?

    OpenAI12 May 2025Leader: Claude Mythos 5 (66.0%)

  • SciCodeLive

    UIUC, Argonne National Laboratory & University of Chicago (Minyang Tian et al.)18 July 2024Leader: OpenAI o3-mini-low (10.8% (main problems))

  • ChemBenchLive

    Jablonka Lab, Friedrich Schiller University Jena1 April 2024Leader: OpenAI o1 (~92% (ChemBench-Mini))

  • MLE-benchLive

    OpenAI9 October 2024

  • LAB-BenchLive

    FutureHouse14 July 2024

  • PaperBenchLive

    OpenAI2 April 2025

  • RE-BenchLive

    METR22 November 2024

  • SciBenchLive

    UCLA, Caltech & University of Washington (Wang, Hu, Lu et al.)20 July 2023

Knowledge & factuality

What a model knows, and whether it admits what it does not.

  • MMLUSaturated

    How much a model knows across 57 academic and professional subjects, tested as four-option multiple-choice questions from elementary to expert level.

    UC Berkeley (Hendrycks et al.)7 September 2020Leader: o1 (91.8%)

  • IFEvalLive

    Google Research14 November 2023

  • FRAMESLive

    Google DeepMind & Harvard University19 September 2024

  • SimpleQALive

    OpenAI30 October 2024Leader: GPT-4.5 (preview) (62.5%)

  • SimpleQA VerifiedLive

    Google DeepMind9 September 2025Leader: Gemini 2.5 Pro (F1 55.6)

  • TruthfulQALive

    Oxford Future of Humanity Institute & OpenAI (Lin, Hilton & Evans)8 September 2021

  • TriviaQASaturating

    University of Washington (Joshi, Choi, Weld & Zettlemoyer)9 May 2017

Long context & retrieval

Finding and using information across very long inputs.

Language & multilingual

Understanding, translation and reasoning beyond English.