Benchmarks · Real-world & economic value

GDPval

also: GDPval-AA

Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.

OpenAIReleased 25 September 2025Live

Most benchmarks ask a model to answer a test question. GDPval asks something closer to what an employer would ask: can the model produce a deliverable — a legal brief, an engineering drawing, a nursing care plan, a customer-support transcript — that a professional in that field would accept as good work? OpenAI built the tasks from 44 occupations across the nine sectors that contribute most to US GDP, each one drafted and checked by an industry professional with an average of 14 years’ experience, so the test items resemble real work products rather than academic exam questions.

Grading is blind: human experts compare a model’s output against an actual human-produced answer to the same task, without being told which is which. At launch, OpenAI reported that GPT-5 and Claude Opus 4.1 — the top two systems it tested — produced work rated as equal to or better than the human comparison on close to half of tasks, and that performance had been improving steadily across model generations. Only a 220-task “gold” subset was released publicly, graded through an automated evaluation service OpenAI hosts, rather than the full 1,320-task set.

GDPval arrived as part of a broader shift, visible the same month in Scale AI’s SWE-bench Pro, away from saturating academic tests and toward harder evaluations grounded in professional output. Its explicit framing around GDP-weighted occupations rather than a research subfield made it a reference point in debates about AI’s effect on knowledge work specifically, and third-party trackers such as Artificial Analysis have since built leaderboards on top of the public gold subset as newer models are released.

The set

1,320 tasks across 44 occupations drawn from the nine US industry sectors that contribute most to GDP, each built and vetted by a professional averaging 14 years' experience; a 220-task 'gold' subset is public, graded via an automated evaluation service OpenAI hosts. Outputs are compared blind against human-produced work on the same task.

Example

An Accountants and Auditors task from the public gold subset: 'You are an auditor and as part of an audit engagement, you are tasked with reviewing and testing the accuracy of reported Anti-Financial Crime Risk Metrics. The attached spreadsheet titled 'Population' contains Anti-Financial Crime Risk Metrics for Q2 and Q3 2024. ... Calculate the required sample size for audit testing based on a 90% confidence level and a 10% tolerable error rate...', supplied with a reference spreadsheet and graded against a human auditor's own deliverable.huggingface.co

Where it stands

Introduced in September 2025 with a win-rate metric (share of tasks rated at least as good as an expert's). A separate third-party GDPval-AA leaderboard (Artificial Analysis) now scores new frontier models on the public gold subset as an Elo rating; on that variant Claude Fable 5 led at 1932 in mid-2026.

Editions, and how each was led

GDPval-AA v2.1current

Released September 2026Artificial Analysis's recalibrated edition (220 tasks, agentic loop with shell and web access). Rates the same models roughly 100–120 Elo below v2, so scores are only comparable within v2.1.

  1. September 2026Claude Opus 5.51846 (Elo v2.1)Independently measured by Artificial Analysis; matches the figure in Anthropic's system card.

Current best: Claude Opus 5.5 — 1846 (GDPval-AA Elo v2.1) Top of the Artificial Analysis v2.1 leaderboard at max effort. Claude Fable 5.1 1735, Claude Opus 5 1708, Grok 4.7 1695, GPT-6 Astra 1542 on the same board.

GDPval-AA v2Retired

Released September 2026A refreshed edition of the Artificial Analysis Elo leaderboard; Elos are not comparable to v1. Superseded within weeks by v2.1, which re-rated the same models roughly 100–120 points lower.

  1. July 2026Claude Opus 51824 (Elo v2)On the GDPval-AA v2 leaderboard, as cited in later launch tables.
  2. September 2026Claude Fable 5.11853 (Elo v2)The current top on v2; Anthropic-reported.

Current best: Claude Fable 5.1 — 1853 (GDPval-AA Elo v2) Anthropic's Fable 5.1 launch table. Opus 5 1824, GPT-5.6 Sol 1710 on comparison tables.

GDPval-AA v1Saturating

Released January 2026Artificial Analysis's third-party Elo over the public gold subset — the version the frontier chart first tracked.

  1. June 2026Claude Fable 5 / Mythos 51932 (Elo v1)Anthropic-reported, on the Artificial Analysis GDPval-AA v1 leaderboard.

Current best: Claude Fable 5 / Mythos 5 — 1932 (GDPval-AA Elo v1) Top of Artificial Analysis's GDPval-AA v1 table. Claude Opus 4.8 1890, GPT-5.5 1769, Gemini 3.1 Pro 1314 on the same table.

OpenAI win-rate (original)

Released 25 September 2025OpenAI's launch metric — a blind win-rate against human deliverables, with no single ranked leaderboard score.

    Current best: GPT-5 / Claude Opus 4.1 — ≈ half of tasks rated ≥ expert The two top systems in OpenAI's own launch evaluation, on the original win-rate metric; OpenAI did not publish a ranked leaderboard score.

    In the timeline · 9 entries

    More real-world & economic value benchmarks