Benchmarks · Knowledge & factuality

MMLU

also: Massive Multitask Language Understanding

How much a model knows across 57 academic and professional subjects, tested as four-option multiple-choice questions from elementary to expert level.

UC Berkeley (Hendrycks et al.)Released 7 September 2020Saturated

For most of the early 2020s, if a new language model was described in a single number, that number was its MMLU score. Dan Hendrycks and colleagues at UC Berkeley published the benchmark in 2020 as a deliberately broad test of knowledge: 15,908 multiple-choice questions across 57 subjects, from elementary arithmetic to professional law and clinical medicine, chosen so that doing well required genuine world knowledge rather than a single linguistic trick.

Its breadth was exactly what made it useful. A model’s average across all 57 subjects compressed a great deal of competence into one comparable figure, and MMLU appeared in essentially every frontier release from GPT-4 onward. That same breadth is what let it saturate: as leading models climbed into the high 80s and low 90s, they approached the accuracy ceiling set by errors in the questions themselves, and the gap between “state of the art” and “as good as the test can measure” closed.

The response was a familiar one in this record — build a harder version. The same group released MMLU-Pro in 2024, with more answer options and more reasoning-heavy questions, and the field’s attention moved on. MMLU’s arc — rapid adoption as the universal yardstick, then rapid saturation — became the template that later benchmarks, from GPQA to Humanity’s Last Exam, were explicitly designed to outrun.

The set

15,908 four-option multiple-choice questions spanning 57 subjects — from elementary mathematics and US history to professional law, clinical medicine and moral reasoning. Scored on plain accuracy; human expert accuracy is estimated at around 90%.

Example

Joe was in charge of lights for a dance. The red light blinks every two seconds, the yellow light every three seconds, and the blue light every five seconds. If we include the very beginning and very end of the dance, how many times during a seven minute dance will all the lights come on at the same time? (Assume that all three lights blink simultaneously at the very beginning of the dance.) A) 3 B) 15 C) 6 D) 5arxiv.org

Where it stands

Frontier models now score in the high 80s and low 90s, near the ceiling set by label noise in the questions themselves; the same authors published the harder MMLU-Pro in 2024, which the field has largely moved to.

How the top score changed hands

  1. September 2020GPT-3 (175B)~43.9%About 20 points above random guessing, far short of expert level.
  2. March 2023GPT-486.4%The result that made MMLU the default headline number for a model release.
  3. June 2024Frontier models (saturated)~88–92%Top models cluster near the noise ceiling; MMLU-Pro is released as the harder successor.

In the timeline · 23 entries · showing 16 most notable

More knowledge & factuality benchmarks