CAIS and Scale AI unveil Humanity's Last Exam results
A 2,500-question expert benchmark built from submissions by nearly 1,000 academics found every frontier model, including o1 and GPT-4o, scored under 10%.
- Benchmarks & progress
- Notable
The Center for AI Safety (CAIS) and Scale AI published Humanity’s Last Exam, a 2,500-question benchmark assembled from submissions by nearly 1,000 contributors across more than 500 institutions in around 50 countries, most of them academics or researchers submitting questions in their own field. The set spans mathematics, the humanities and the natural sciences in both text-only and multimodal formats, filtered from a much larger pool of submitted questions for ones judged unanswerable by internet search and gradable automatically.
The benchmark was designed explicitly against saturation. CAIS director Dan Hendrycks argued that existing tests such as MMLU had stopped being informative once models cleared roughly 90% accuracy, leaving no way to measure further progress at the frontier. At launch, CAIS and Scale AI reported that every model tested — including OpenAI’s o1 and GPT-4o — answered fewer than 10% of questions correctly, with the models often as confident in wrong answers as in right ones.
The low starting scores were the intended design rather than a criticism of the models; Hendrycks said the team could not predict how quickly scores would rise, which was itself the point of building the test. They did rise substantially: within about a year, leading models on Scale AI’s public leaderboard scored above 40%, a large jump from the low single digits at launch though still short of saturation. Humanity’s Last Exam joined a run of benchmarks published in this period — including FrontierMath — built specifically to outlast the rapid gains that had exhausted their predecessors.