Benchmarks · Reasoning & problem-solving

EnigmaEval

also: Enigma Eval, EnigmaEval benchmark

Long-horizon multimodal reasoning: synthesising implicit knowledge and chaining many steps of lateral deduction over mixed image-and-text puzzle-hunt problems that take skilled human teams hours to days to solve.

Scale AI & Center for AI SafetyReleased 13 February 2025Live

EnigmaEval is built out of puzzle hunts — the elaborate, multi-layered puzzles that teams of enthusiasts spend hours or days unpicking at events like the MIT Mystery Hunt. Scale AI and the Center for AI Safety transcribed 1,184 of them into a mixed image-and-text form, each with a single verifiable answer, precisely because such puzzles resist the shortcuts that saturate ordinary benchmarks: there is no template to pattern-match, the relevant knowledge is implicit, and the path from clue to answer runs through several leaps of lateral reasoning.

That design makes EnigmaEval one of the hardest tests in circulation. A puzzle like “Switching Channels” hands the solver two columns of television shows and asks them not merely to pair the shows but to discover the hidden theme those pairings encode — the kind of step that is obvious in hindsight and very hard to reach. The benchmark keeps a Normal split and a Hard split drawn from the most demanding hunts, so a single number still separates strong reasoners from the rest.

EnigmaEval is tracked on the Center for AI Safety’s AI Dashboard as an independent third-party evaluation, and there it remains among the least saturated: GPT-4o solved well under 1% of puzzles, reasoning models such as o3 lifted that into the teens, and by mid-2026 the frontier — Claude Opus 5 at 43.9% — had reached the low 40s while still failing the majority. Its slow climb is the reason it has become a standard citation for reasoning that unfolds over many steps rather than one.

The set

1,184 puzzles drawn from real-world puzzle hunts and transcribed by human annotators into a text-and-image form, each with a single unambiguous answer graded by string match. A Normal split (949 puzzles, from events such as PuzzledPint and cryptic crosswords) and a much harder Hard split (235, from the MIT Mystery Hunt and similar) separate difficulty.

Example

'Switching Channels', a Normal-split puzzle: solvers are given two columns of television shows and told the first show is featured in the left column and its partner is somewhere in the right column; they must pair the shows and then work out the hidden theme that unifies each pairing — the pairing is not the answer, only the route to it.arxiv.org

Where it stands

One of the least-saturated evaluations on the independent CAIS AI Dashboard: frontier models climbed from under 1% in 2024 to the low 40s by mid-2026, still leaving most puzzles unsolved.

How the top score changed hands

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. May 2024GPT-4o0.8%Effectively unable to solve the puzzles at all on the CAIS AI Dashboard.
  2. December 2024o311.9%Reasoning models moved the score off the floor.
  3. November 2025Gemini 3 Pro17.8%
  4. February 2026Gemini 3.1 Pro32.4%A large jump, past the 30% mark.
  5. June 2026Claude Fable 541.3%
  6. July 2026Claude Opus 543.9%The highest EnigmaEval score recorded on the CAIS dashboard here.

Current best: Claude Opus 5 — 43.9% On the independent CAIS AI Dashboard, which runs EnigmaEval as a standardized third-party evaluation across frontier models.

More reasoning & problem-solving benchmarks