Reference
AI benchmarks
A benchmark is a fixed test that lets different AI systems be compared on the same footing. This catalogue is ordered by how live each one is: the benchmarks the frontier labs actively compete on — many scores, from many labs, with the lead still changing hands — sit largest and first, marked 🔥 Competitive; the ones that were an interesting idea but never grew a live leaderboard trail off smaller below. The recurring cycle — a benchmark is built to be hard, models climb it within a year or two, it saturates, a harder one replaces it — is followed as a story in the benchmarks and their saturation thread.
118 benchmarks · 12 domains · 32 competitive
Coding & software engineering
Writing code, fixing real bugs, resolving issues in live repositories.
DeepSWE🔥 CompetitiveLive
Can a coding agent complete an original, long-horizon software engineering task in a real repository, graded by whether the behaviour is correct — not just whether it matches one specific reference implementation?
SWE-bench Pro🔥 CompetitiveLive
Can a model resolve a realistic, multi-file software engineering task in a codebase it could not have memorised — including private, commercial code rather than only well-known open-source repositories?
SWE-bench🔥 Competitive
Can a model resolve a real, unseen GitHub issue by editing a codebase so that the project's own hidden tests pass?
Codeforces / CodeContests🔥 CompetitiveLive
Can a model solve genuinely novel algorithmic problems under contest conditions — the kind that require devising an approach, not recalling one — well enough to rank against real competitive programmers?
Terminal-Bench🔥 CompetitiveLive
Can an AI agent actually operate a computer through a real command-line shell — issuing commands, reading their output, and adapting — to finish a multi-step task, rather than just producing plausible-looking commands?
LiveCodeBenchLive
How well a model codes on problems it could not have memorised, by dating every problem and checking performance separately on those published before and after the model's training cutoff.
HumanEvalSaturated
Can a model write a correct, working Python function from a natural-language docstring alone?
Aider PolyglotLive
Can a model act as a practical pair-programmer — reading an existing multi-language codebase, understanding a task, and editing the actual files correctly, not just writing an isolated function?
SWE-LancerLive
NanoGPT Speedrun🔥 CompetitiveLive
MBPPSaturated
BigCodeBenchLive
Konwinski PrizeRetired
Reasoning & problem-solving
General reasoning, hard exams, puzzles that resist memorisation.
GPQA🔥 Competitive
Whether a model can answer graduate-level science questions that a skilled non-expert cannot solve even with unrestricted web access and half an hour per question.
Humanity's Last Exam🔥 CompetitiveLive
Whether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.
EnigmaEval🔥 CompetitiveLive
Long-horizon multimodal reasoning: synthesising implicit knowledge and chaining many steps of lateral deduction over mixed image-and-text puzzle-hunt problems that take skilled human teams hours to days to solve.
ARC-AGI🔥 CompetitiveLive
Whether a system can infer an unfamiliar abstract rule from a handful of examples and apply it to a new case, rather than recognising a pattern it has seen before.
MMLU-Pro🔥 CompetitiveLive
Whether a model's knowledge holds up once guessing is made hard and questions demand multi-step reasoning rather than recall.
SimpleBenchLive
DROPRetired
BIG-BenchRetired
AGIEvalSaturated
HellaSwagSaturated
WinoGrandeRetired
Multimodal
Vision, audio, video and charts — reasoning over more than text.
MMMU🔥 CompetitiveLive
Can a model answer college-exam-level questions that genuinely require reading an accompanying image — a chart, diagram, map or chemical structure — rather than knowledge alone?
MindCube🔥 CompetitiveLive
Whether a vision-language model can build a spatial mental model of a scene — inferring the positions, orientations and possible movements of objects, including ones it cannot currently see — from only a few limited views.
SpatialViz🔥 CompetitiveLive
Spatial visualization: whether a multimodal model can mentally imagine and manipulate visual structures that are not directly observable, across mental rotation, mental folding, visual penetration and mental animation.
ERQA🔥 CompetitiveLive
Embodied Reasoning QA: whether a vision-language model understands a physical scene well enough to reason about acting in it — spatial relations, trajectories, state estimation, pointing and multi-view correspondence.
IntPhys 2🔥 CompetitiveLive
Intuitive physics from video: whether a model grasps four macroscopic principles — object permanence, immutability, spatio-temporal continuity and solidity — well enough to tell physically possible events from impossible ones.
Video-MMELive
BLINKLive
RealWorldQALive
MVBenchLive
DocVQASaturated
VQAv2Saturated
Agents, tools & computer use
Multi-step tasks: driving a browser, a terminal, a desktop, real tools.
AutomationBench🔥 CompetitiveLive
Can an AI agent carry out a real business workflow across multiple SaaS applications — finding the right API endpoints itself, following a company's own layered business rules, and getting the right data to the right system — rather than completing a single, well-specified task?
TextQuests🔥 CompetitiveLive
Long-horizon agentic reasoning: whether an LLM can autonomously play classic text-adventure games, sustaining exploration, planning and memory over a long and growing context without external tools.
OSWorld🔥 CompetitiveLive
Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?
BrowseComp🔥 CompetitiveDisputed
Can an agent find a specific, hard-to-locate fact on the open web by searching persistently and connecting scattered clues, rather than by knowing the answer already or finding it in one search?
τ-benchLive
GAIALive
VisualWebArenaLive
WebVoyagerLive
TheAgentCompanyLive
AgentBenchRetired
WebArenaLive
Online-Mind2WebDisputed
Aggregate indices & arenas
Human-preference arenas and composite indices that roll many tests into one ranking.
LMArena🔥 CompetitiveDisputed
Which of two anonymous models people prefer in open-ended, head-to-head conversation, aggregated into a running Elo rating rather than a fixed test score.
Artificial Analysis Intelligence Index🔥 CompetitiveLive
A single composite score summarising a model's capability across agentic tasks, coding, scientific reasoning and general knowledge, built as a weighted average of several independent evaluations Artificial Analysis runs itself rather than reported by the model developer.
LiveBenchLive
How a model performs on recently created questions, graded by objective ground-truth answers rather than human or LLM judgment, so that scores cannot reflect memorised test data and cannot be inflated by a biased judge model.
Artificial Analysis Coding Agent Index🔥 CompetitiveLive
How well a coding agent — a specific model paired with a specific harness, such as Claude Code or Codex — completes real software-engineering work end to end, not just whether the underlying model answers a coding question correctly.
HELMLive
Open LLM LeaderboardRetired
Mathematics
From grade-school word problems to unsolved research mathematics.
AIME🔥 CompetitiveContamination concerns
Whether a model can solve short, single-answer competition-maths problems requiring several steps of reasoning but no written proof.
MATH🔥 CompetitiveSaturated
Can a model solve a competition-level mathematics problem and produce a correct step-by-step derivation, not just a lucky final number?
FrontierMath🔥 CompetitiveDisputed
Can a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?
GSM8KSaturated
Can a model solve a grade-school arithmetic word problem that takes several linked steps to work through, rather than a single calculation?
MathArenaLive
miniF2FSaturated
HARPLive
PutnamBenchLive
HMMTLive
Omni-MATHLive
Real-world & economic value
Professional work products and economically valuable tasks.
GDPval🔥 CompetitiveLive
Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.
Agents' Last Exam🔥 CompetitiveLive
Whether an AI agent can complete long-horizon, economically valuable professional tasks — not just answer questions — with a verifiable, checkable outcome.
Vending-BenchLive
WorkBenchLive
CORE-BenchLive
Safety, security & robustness
Jailbreaks, dangerous-capability evals, honesty, sabotage, harm.
ExploitBench🔥 CompetitiveLive
How far an AI agent gets through the actual chain of an exploit — not just whether it crashes a target, but whether it can turn that crash into control of the machine.
ExploitGym🔥 CompetitiveLive
Can an AI agent turn a known software vulnerability into a real, working attack — not merely identify or patch it, but exploit it end to end, including against active defences?
SEC-bench Pro🔥 CompetitiveLive
Can a model find a genuine, previously undisclosed-style vulnerability in a large, real codebase and prove it with a working exploit input — not just patch a bug it has already been shown?
MASKLive
CyberGymLive
AgentHarmLive
SHADE-ArenaLive
HarmBenchLive
StrongREJECTLive
CybenchLive
CyberSecEvalLive
WMDPLive
JailbreakBenchLive
SEC-benchLive
Science & research
Domain knowledge and research work in the sciences.
HealthBench🔥 CompetitiveLive
How well does a model handle realistic, open-ended health conversations — with a layperson or a clinician — judged against criteria that practising physicians say actually matter, rather than a multiple-choice medical exam?
SciCodeLive
ChemBenchLive
MLE-benchLive
LAB-BenchLive
PaperBenchLive
RE-BenchLive
SciBenchLive
Knowledge & factuality
What a model knows, and whether it admits what it does not.
MMLUSaturated
How much a model knows across 57 academic and professional subjects, tested as four-option multiple-choice questions from elementary to expert level.
IFEvalLive
FRAMESLive
SimpleQALive
TruthfulQALive
Long context & retrieval
Finding and using information across very long inputs.
LongBenchLive
RULERLive
MRCRLive
Needle in a HaystackSaturated
Language & multilingual
Understanding, translation and reasoning beyond English.
CMMLULive
MGSMLive
BelebeleLive
FLORES-200Live
C-EvalRetired
Global-MMLULive
IndQALive
MMMLULive