Benchmarks · Knowledge & factuality

SimpleQA

also: Simple QA

Whether a model gives a correct, confidently stated answer to a short factual question that has exactly one verified answer, or appropriately admits it doesn't know.

OpenAIReleased 30 October 2024Live

Most factuality benchmarks mix reasoning, comprehension and recall together, which makes it hard to tell whether a wrong answer came from bad knowledge or bad reasoning. OpenAI built SimpleQA to isolate the first problem: 4,326 short questions, each adversarially chosen to have one verified, unambiguous answer, so that a model’s response can only be graded correct, incorrect, or — crucially — not attempted. That third category was the point. Rather than only rewarding accuracy, the benchmark also credits a model for declining to answer when it is unsure, treating a confident wrong answer as worse than an honest “I don’t know.”

At launch, no model came close to acing it: OpenAI’s own o1-preview led with 42.7% correct, ahead of GPT-4o’s 38.2%, while contemporary Claude models scored under 30%. The gap closed quickly. By the time GPT-4.5 shipped in February 2025, OpenAI’s results table put it at 62.5% — the company framed the same improvement as a drop in confabulation rate, from roughly 60% for GPT-4o to 37% for GPT-4.5.

SimpleQA became a standard reference for hallucination and calibration rather than raw capability, cited across OpenAI, Google, xAI and DeepSeek model cards through 2025 — though notably not Anthropic’s, which reports factuality through other evals. It also produced a counter-intuitive result that has held: GPT-4.5, a large non-reasoning model tuned for breadth of knowledge, still tops the closed-book board at 62.5%, ahead of the reasoning models that followed it — o3 around 49%, Gemini 2.5 Pro 54%, GPT-5 thinking 55% — which trade some factual recall for reasoning depth.

Two things complicate reading a “SimpleQA” number. The score is usually reported as overall % correct, paired with a hallucination rate (incorrect ÷ attempted) — a different figure that some tables show alongside it. And in September 2025 Google DeepMind released SimpleQA-Verified, a cleaned 1,000-question subset scored on an F1 metric; it shares the name and heritage but not the scale, and its numbers must be kept separate. OpenAI’s simple-evals repository stopped adding new model results in July 2025, so later frontier scores live in individual model cards rather than one continuously updated table — and no 2026 flagship has a confirmed closed-book original-SimpleQA figure.

The set

4,326 short, adversarially collected fact-seeking questions spanning history, science, technology, art and more, each checked by two independent reviewers for a single unambiguous answer. Responses are graded correct, incorrect, or not attempted, since the authors wanted the benchmark to reward calibration rather than confident guessing.

Example

'Who received the IEEE Frank Rosenblatt Award in 2010?' (Answer: Michio Sugeno) — one of the paper's own worked examples of a single-answer, adversarially checked factual question.arxiv.org

Where it stands

GPT-4.5's 62.5% (Feb 2025) remains the highest confirmed closed-book score on the original SimpleQA — a telling result, since the reasoning-focused models that followed (o3 ~49%, Gemini 2.5 Pro 54%, GPT-5 thinking 55%) score lower on pure recall. OpenAI's simple-evals repo stopped adding models in July 2025, and no 2026 flagship has a confirmed closed-book original-SimpleQA figure. Don't confuse it with SimpleQA-Verified, a separate cleaned 1,000-question Google subset scored on a different (F1) metric.

How the top score changed hands

  1. October 2024OpenAI o1-preview42.7%Best-scoring model in the launch paper's own results table; GPT-4o scored 38.2% and contemporary Claude models under 30%.
  2. February 2025GPT-4.5 (preview)62.5%Reported alongside OpenAI's own hallucination-rate framing of the same evaluation.

Current best: GPT-4.5 (preview) — 62.5% The highest confirmed % correct on the original (closed-book) SimpleQA — and it has held, because the big non-reasoning GPT-4.5 was tuned for factual breadth, whereas the reasoning models after it (o3, Gemini 2.5 Pro, GPT-5) score lower on pure recall (~49–55%). Metric is overall % correct; model cards pair it with a hallucination rate (incorrect ÷ attempted). Anthropic does not report SimpleQA.

In the timeline · 2 entries

More knowledge & factuality benchmarks