Benchmarks · Reasoning & problem-solving

Humanity's Last Exam

also: HLE

Whether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.

Center for AI Safety & Scale AIReleased 23 January 2025Live

Humanity’s Last Exam was built to solve a specific problem: MMLU and its peers had stopped being useful once frontier models cleared roughly 90% accuracy, leaving no way to see continued progress. CAIS and Scale AI assembled 2,500 questions submitted by nearly 1,000 academics across roughly 500 institutions, each one filtered to be unanswerable by internet search and gradable by a fixed answer key. At launch in January 2025, the design succeeded almost too well — every model tested, including OpenAI’s o1 and GPT-4o, scored under 10%, and CAIS director Dan Hendrycks said the team genuinely did not know how fast scores would climb.

They climbed fast. Reasoning-focused models pushed past 40% within the year: xAI reported 44.4% for a multi-agent configuration of Grok 4, though the figure had not yet cleared independent verification on the public leaderboard, and Gemini 3 Pro reached 37.5% without tools, rising to 41.0% in its higher-effort Deep Think mode. By February 2026, an upgraded Deep Think reported 48.4% without external tools, and by mid-2026 the independent CAIS AI Dashboard had Claude Fable 5 topping its no-tools leaderboard at 52.7%; Anthropic separately reported a higher 59.0% for the larger Claude Mythos 5.

Scores with tool access — letting a model search or use a code interpreter mid-answer — have run higher still, which makes cross-model comparison harder: a number is only meaningful alongside the conditions it was measured under. HLE has nonetheless become, alongside GPQA and ARC-AGI, one of the benchmarks labs reach for specifically because it has not yet saturated, and its scores remain a standard citation in frontier model announcements.

The set

2,500 questions submitted by nearly 1,000 contributors across roughly 500 institutions in about 50 countries, spanning mathematics, the humanities and the natural sciences in both text-only and multimodal form. Questions were filtered to exclude anything answerable by internet search and to remain automatically gradable.

Example

Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.lastexam.ai

Where it stands

Every model tested at launch in January 2025 scored under 10%; by mid-2026 the independent CAIS dashboard put the no-tools leader (Claude Fable 5) at 52.7%, with Anthropic separately reporting Claude Mythos 5 at 59.0% and, in September 2026, Claude Fable 5.1 at 60.9% (no tools) — still well short of saturation and not independently verified.

Editions, and how each was led

HLE (full, no tools)current

Released 23 January 2025The original 2,500-question set, scored without external tools — the condition the frontier chart tracks.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. January 2025o1 / GPT-4ounder 10%Every model CAIS and Scale AI tested at launch, across the board, scored below 10%.
  2. July 2025Grok 4 Heavy44.4%Multi-agent 'Heavy' tier; xAI's figure had not yet appeared on the public leaderboard for independent verification.
  3. November 2025Gemini 3 Pro37.5% (41.0% Deep Think)Without external tools.
  4. February 2026Gemini 3 Deep Think (v2)48.4%Without external tools.
  5. June 2026Claude Fable 552.7%No tools; tops the independent CAIS AI Dashboard no-tools leaderboard. Anthropic separately reported Claude Mythos 5 at 59.0%.

Current best: Claude Fable 5 — 52.7% No tools, on the independent CAIS AI Dashboard, which tops the no-tools leaderboard. Anthropic separately reported higher no-tools figures in its own launch tables — 59.0% for the larger Claude Mythos 5, and 60.9% for Claude Fable 5.1 in September 2026 — and Moonshot AI reported 54.0% for Kimi K2.6 with tool use, a different, tool-assisted condition; none of these vendor figures are independently verified.

HLE-Verified

Released August 2026A cleaned 'Verified' subset that removes questions with errors or ambiguous answers, cited by the newest launches; not comparable to the full set.

  1. June 2026GPT-5.6 Sol54.5%On HLE-Verified, as listed in Google's Gemini 3.8 Flash comparison table.
  2. September 2026Gemini 3.8 Flash54.9%The top figure on Google's HLE-Verified comparison; company-reported.

Current best: Gemini 3.8 Flash — 54.9% Top of Google's Gemini 3.8 Flash comparison table on HLE-Verified; GPT-5.6 Sol 54.5%, Claude Opus 5 54.4% on the same table. Company-reported.

In the timeline · 14 entries

More reasoning & problem-solving benchmarks