Benchmarks · Real-world & economic value
GDPval
also: GDPval-AA
Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.
OpenAIReleased 25 September 2025Live
Most benchmarks ask a model to answer a test question. GDPval asks something closer to what an employer would ask: can the model produce a deliverable — a legal brief, an engineering drawing, a nursing care plan, a customer-support transcript — that a professional in that field would accept as good work? OpenAI built the tasks from 44 occupations across the nine sectors that contribute most to US GDP, each one drafted and checked by an industry professional with an average of 14 years’ experience, so the test items resemble real work products rather than academic exam questions.
Grading is blind: human experts compare a model’s output against an actual human-produced answer to the same task, without being told which is which. At launch, OpenAI reported that GPT-5 and Claude Opus 4.1 — the top two systems it tested — produced work rated as equal to or better than the human comparison on close to half of tasks, and that performance had been improving steadily across model generations. Only a 220-task “gold” subset was released publicly, graded through an automated evaluation service OpenAI hosts, rather than the full 1,320-task set.
GDPval arrived as part of a broader shift, visible the same month in Scale AI’s SWE-bench Pro, away from saturating academic tests and toward harder evaluations grounded in professional output. Its explicit framing around GDP-weighted occupations rather than a research subfield made it a reference point in debates about AI’s effect on knowledge work specifically, and third-party trackers such as Artificial Analysis have since built leaderboards on top of the public gold subset as newer models are released.
The set
1,320 tasks across 44 occupations drawn from the nine US industry sectors that contribute most to GDP, each built and vetted by a professional averaging 14 years' experience; a 220-task 'gold' subset is public, graded via an automated evaluation service OpenAI hosts. Outputs are compared blind against human-produced work on the same task.
Example
An Accountants and Auditors task from the public gold subset: 'You are an auditor and as part of an audit engagement, you are tasked with reviewing and testing the accuracy of reported Anti-Financial Crime Risk Metrics. The attached spreadsheet titled 'Population' contains Anti-Financial Crime Risk Metrics for Q2 and Q3 2024. ... Calculate the required sample size for audit testing based on a 90% confidence level and a 10% tolerable error rate...', supplied with a reference spreadsheet and graded against a human auditor's own deliverable.huggingface.co
Where it stands
Introduced in September 2025 with a win-rate metric (share of tasks rated at least as good as an expert's). A separate third-party GDPval-AA leaderboard (Artificial Analysis) now scores new frontier models on the public gold subset as an Elo rating; on that variant Claude Fable 5 led at 1932 in mid-2026.
Editions, and how each was led
GDPval-AA v2.1current
Current best: Claude Opus 5.5 — 1846 (GDPval-AA Elo v2.1) Top of the Artificial Analysis v2.1 leaderboard at max effort. Claude Fable 5.1 1735, Claude Opus 5 1708, Grok 4.7 1695, GPT-6 Astra 1542 on the same board.
GDPval-AA v2Retired
- July 2026Claude Opus 51824 (Elo v2)On the GDPval-AA v2 leaderboard, as cited in later launch tables.
- September 2026Claude Fable 5.11853 (Elo v2)The current top on v2; Anthropic-reported.
Current best: Claude Fable 5.1 — 1853 (GDPval-AA Elo v2) Anthropic's Fable 5.1 launch table. Opus 5 1824, GPT-5.6 Sol 1710 on comparison tables.
GDPval-AA v1
Current best: Claude Fable 5 / Mythos 5 — 1932 (GDPval-AA Elo v1) Top of Artificial Analysis's GDPval-AA v1 table. Claude Opus 4.8 1890, GPT-5.5 1769, Gemini 3.1 Pro 1314 on the same table.
OpenAI win-rate (original)
Current best: GPT-5 / Claude Opus 4.1 — ≈ half of tasks rated ≥ expert The two top systems in OpenAI's own launch evaluation, on the original win-rate metric; OpenAI did not publish a ranked leaderboard score.
In the timeline · 9 entries
Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
Models & capabilities · Benchmarks & progress
xAI releases Grok 4.7
xAI reported a larger base model than Grok 4.6, up to 2.1 trillion parameters against a reported 1.5 trillion, at unchanged pricing; independent ranking placed it sixth on Artificial Analysis's GDPval-AA index.
Models & capabilities
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1
Anthropic's launch table put Fable 5.1 ahead of its own Fable 5 and Opus 5 and OpenAI's GPT-5.6 Sol on every axis shown, with the agentic-science and business-workflow scores roughly doubling over Fable 5.
Models & capabilities · Benchmarks & progress
xAI releases a new Grok model that a benchmark firm rates on par with OpenAI's flagship
Priced the same as its predecessor at $2/$6 per million tokens; Musk said a larger Grok 4.7 was already in training and expected within three to four weeks.
Models & capabilities
ByteDance releases Seed 2.1 model family
The closed-weight Pro and Turbo models target multi-step agent tasks such as mobile-app control and coding, sold through Volcano Engine at prices quoted in yuan.
Models & capabilities
UC Berkeley releases Agents' Last Exam, a benchmark of professional work
Built with 250+ industry experts across 55 sub-industries, it runs agents in the real software a specialist would use and grades against hidden answers; current systems clear under 1% of the hardest tier.
Benchmarks & progress
OpenAI releases GPT-5.2
Released three weeks after Google's Gemini 3 and following a reported internal OpenAI 'code red,' with a claimed 70.9% win rate against professionals on the GDPval benchmark, up from 38.8% for GPT-5.1.
Models & capabilities
OpenAI publishes GDPval, a benchmark for economically valuable knowledge work
Blind grading by industry professionals rated GPT-5 and Claude Opus 4.1 outputs as equal to or better than human work on nearly half of the 1,320 tasks.
Benchmarks & progress