Benchmarks · Knowledge & factuality
MMLU
also: Massive Multitask Language Understanding
How much a model knows across 57 academic and professional subjects, tested as four-option multiple-choice questions from elementary to expert level.
UC Berkeley (Hendrycks et al.)Released 7 September 2020Saturated
For most of the early 2020s, if a new language model was described in a single number, that number was its MMLU score. Dan Hendrycks and colleagues at UC Berkeley published the benchmark in 2020 as a deliberately broad test of knowledge: 15,908 multiple-choice questions across 57 subjects, from elementary arithmetic to professional law and clinical medicine, chosen so that doing well required genuine world knowledge rather than a single linguistic trick.
Its breadth was exactly what made it useful. A model’s average across all 57 subjects compressed a great deal of competence into one comparable figure, and MMLU appeared in essentially every frontier release from GPT-4 onward. That same breadth is what let it saturate: as leading models climbed into the high 80s and low 90s, they approached the accuracy ceiling set by errors in the questions themselves, and the gap between “state of the art” and “as good as the test can measure” closed.
The response was a familiar one in this record — build a harder version. The same group released MMLU-Pro in 2024, with more answer options and more reasoning-heavy questions, and the field’s attention moved on. MMLU’s arc — rapid adoption as the universal yardstick, then rapid saturation — became the template that later benchmarks, from GPQA to Humanity’s Last Exam, were explicitly designed to outrun.
The set
15,908 four-option multiple-choice questions spanning 57 subjects — from elementary mathematics and US history to professional law, clinical medicine and moral reasoning. Scored on plain accuracy; human expert accuracy is estimated at around 90%.
Example
Joe was in charge of lights for a dance. The red light blinks every two seconds, the yellow light every three seconds, and the blue light every five seconds. If we include the very beginning and very end of the dance, how many times during a seven minute dance will all the lights come on at the same time? (Assume that all three lights blink simultaneously at the very beginning of the dance.) A) 3 B) 15 C) 6 D) 5arxiv.org
Where it stands
Frontier models now score in the high 80s and low 90s, near the ceiling set by label noise in the questions themselves; the same authors published the harder MMLU-Pro in 2024, which the field has largely moved to.
How the top score changed hands
- September 2020GPT-3 (175B)~43.9%About 20 points above random guessing, far short of expert level.
- March 2023GPT-486.4%The result that made MMLU the default headline number for a model release.
- June 2024Frontier models (saturated)~88–92%Top models cluster near the noise ceiling; MMLU-Pro is released as the harder successor.
In the timeline · 23 entries · showing 16 most notable
OpenAI publishes open-weight models for the first time since GPT-2
gpt-oss-120b runs on a single 80GB GPU and matches OpenAI's own o4-mini on core reasoning benchmarks; the smaller 20b model runs on 16GB of memory.
Open weights & ecosystem
CAIS and Scale AI unveil Humanity's Last Exam results
A 2,500-question expert benchmark built from submissions by nearly 1,000 academics found every frontier model, including o1 and GPT-4o, scored under 10%.
Benchmarks & progress
Epoch AI launches FrontierMath
Built with over 60 mathematicians including Fields medallists as reviewers, the benchmark held leading models under 2% accuracy even with extended reasoning time and code tools.
Benchmarks & progress
Tencent open-sources Hunyuan-Large MoE model
Tencent said the 389B-parameter, 52B-active MoE model beat Llama 3.1 405B on MMLU and MATH despite far fewer active parameters, and released a technical report alongside the weights.
Open weights & ecosystem · Models & capabilities
Alibaba releases Qwen2.5 model family
Alibaba's release spanned seven sizes from 0.5B to 72B parameters, plus dedicated coding and maths variants, trained on 18 trillion tokens.
Open weights & ecosystem · Models & capabilities
xAI releases Grok-2
The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
Models & capabilities
Mistral AI releases Mistral Large 2
The 123-billion-parameter model reported 84.0% on MMLU and was released under a non-commercial research licence, with a separate paid licence for commercial use.
Models & capabilities
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
MMLU-Pro benchmark paper released
The paper reported chain-of-thought reasoning helped on the new benchmark where it had made little difference on the original MMLU, and cut prompt-sensitivity from 4-5 points to about 2.
Benchmarks & progress
Anthropic's Claude 3 takes the frontier from GPT-4
The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
Models & capabilities · Labs & people
Google renames Bard to Gemini and launches Gemini Advanced with Ultra 1.0
The $19.99-a-month Google One AI Premium tier gave access to Ultra 1.0, which Google said was the first model to outperform human experts on MMLU.
Models & capabilities
Google launches Gemini
Google said Gemini Ultra beat human experts on the MMLU benchmark; days later Bloomberg reported the model's showcase video had been edited and was not real-time.
Models & capabilities · Culture & impact
LMSYS launches Chatbot Arena
The Berkeley-linked project ranked chatbots by anonymous, randomised head-to-head votes rather than fixed test sets, and later became LMArena.
Benchmarks & progress
OpenAI releases GPT-4
A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
Models & capabilities · Benchmarks & progress
DeepMind's Chinchilla paper rewrites the scaling laws
Existing large models were badly under-trained: for a fixed compute budget, parameters and training tokens should scale together.
Ideas & essays · Benchmarks & progress · Models & capabilities
Hendrycks et al. publish the MMLU benchmark
15,908-question, 57-subject multiple-choice benchmark spanning elementary to professional level; GPT-3 improved on random chance by roughly 20 points on average.
Benchmarks & progress