Benchmarks · Reasoning & problem-solving
SimpleBench
also: Simple Bench
Whether a model can handle everyday spatio-temporal reasoning, social intelligence and 'trick question' style linguistic traps that an ordinary person finds easy but that memorised knowledge doesn't help with.
SimpleBench TeamReleased October 2024Live
Most benchmarks on this list were built to be hard for reasons that need explaining — graduate-level science, adversarially filtered distractors, thousands of crowd-sourced exam questions. SimpleBench’s premise is the opposite: its questions are supposed to be easy, the kind of everyday spatial, social and common-sense reasoning a person with no specialist training answers correctly without much thought. The benchmark’s more than 200 multiple-choice questions test spatio-temporal reasoning, social intelligence and what its creators call “linguistic adversarial robustness” — trick-question phrasing designed to catch a system relying on surface pattern-matching rather than actually working through the scenario.
That framing produced an unusual result for a benchmark introduced in late 2024: models that were already scoring in the 80s and 90s on graduate-level exams struggled with questions a non-specialist human finds routine. Early models such as GPT-4 Turbo and OpenAI’s o1-preview scored around a quarter and two-fifths respectively, against a human baseline of 83.7% drawn from a small sample of nine participants (with the single highest human scorer reaching 95.4%).
Scores have climbed substantially since — reasoning-focused models from Anthropic, Google, OpenAI and others have pushed well past 70% — but as of SimpleBench’s live leaderboard in August 2026, the best-scoring model, Anthropic’s Claude Fable at 81.9%, still trails the human baseline. That makes SimpleBench one of a small number of benchmarks where “beating the average human” remains, on the benchmark’s own numbers, an open question rather than a solved one.
The set
Over 200 multiple-choice questions covering spatio-temporal reasoning, social intelligence and linguistic adversarial robustness — questions written to resist the pattern-matching and recalled-knowledge strategies that let language models do well on most academic benchmarks.
Example
Beth places four whole ice cubes in a frying pan at the start of the first minute, then five at the start of the second minute and some more at the start of the third minute, but none in the fourth minute. If the average number of ice cubes per minute placed in the pan while it was frying a crispy egg was five, how many whole ice cubes can be found in the pan at the end of the third minute? A) 30 B) 0 C) 20 D) 10 E) 11 F) 5 (Answer: B)github.com
Where it stands
As of the site's live leaderboard checked in August 2026, no model has yet beaten the non-specialist human baseline, though the gap has narrowed sharply since the benchmark launched.
How the top score changed hands
- November 2023GPT-4 Turbo25.1%
- September 2024o1-preview41.7%
- August 2026Claude Fable81.9%Current top score on the live leaderboard, checked August 2026; still short of the 83.7% human baseline.
Current best: Claude Fable — 81.9% Per SimpleBench's live leaderboard, checked August 2026. Still below the 83.7% human baseline (and the 95.4% highest individual human score) measured from a small sample of nine participants.