Benchmarks · Reasoning & problem-solving

MMLU-Pro

also: MMLU Pro

Whether a model's knowledge holds up once guessing is made hard and questions demand multi-step reasoning rather than recall.

TIGER-AI-Lab (Wang, Ma, Zhang et al.)Released 3 June 2024Live

MMLU had a guessing problem: with four answer choices, a model that knew nothing could still score 25%, and by 2024 leading systems were closing in on the accuracy ceiling anyway. TIGER-AI-Lab’s response, published in June 2024, kept MMLU’s format — multiple-choice questions across academic and professional subjects — but expanded each question to ten options and skewed the question mix toward problems that need several reasoning steps rather than a single recalled fact.

The redesign changed which techniques helped. On the original MMLU, chain-of-thought prompting made little difference over a model simply answering directly; on MMLU-Pro it produced a real gain, evidence that the harder questions were rewarding actual reasoning rather than pattern-matched recall. Scores also became more stable: sensitivity to how a question was phrased fell from a 4–5 point swing on MMLU to about 2 points here. At launch, the best-performing model, GPT-4o, scored around 72.6%, some 14 points below its MMLU number, and Claude 3.5 Sonnet moved ahead within weeks at just over 76%.

MMLU-Pro has proved more durable than its predecessor but is not immune to the same pressure: by 2026 the top model on TIGER-Lab’s own leaderboard, Gemini 3.1 Pro, was scoring above 91%, and the gap between frontier systems had narrowed again. It remains one of the standard secondary citations alongside MMLU in model release reports, particularly for labs wanting to show a harder knowledge test than the original still supports.

The set

Over 12,000 curated questions across 14 subjects — biology, business, chemistry, computer science, economics, engineering, health, history, law, math, philosophy, physics, psychology and 'other' — each with ten answer options instead of the original MMLU's four, cutting the odds of a correct guess.

Example

A refracting telescope consists of two converging lenses separated by 100 cm. The eye-piece lens has a focal length of 20 cm. The angular magnification of the telescope is A) 10 B) 40 C) 6 D) 25 E) 15 F) 50 G) 30 H) 4 I) 5 J) 20 (Answer: H)arxiv.org

Where it stands

Frontier models now score above 90%, and chain-of-thought prompting — which barely helped on the original MMLU — produces a clear gain here, one of the design features meant to keep the test informative for longer.

How the top score changed hands

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2024GPT-4o72.55%Best model reported around the benchmark's release.
  2. June 2024Claude 3.5 Sonnet76.12%
  3. February 2026Gemini 3.1 Pro91.16%Current leader on TIGER-Lab's leaderboard, checked August 2026.

Current best: Gemini 3.1 Pro — 91.16% Per TIGER-Lab's own MMLU-Pro leaderboard, checked August 2026; later frontier models may score higher but are not yet listed there.

In the timeline · 5 entries

More reasoning & problem-solving benchmarks