Benchmarks · Real-world & economic value
GDPval
also: GDPval-AA
Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.
OpenAIReleased 25 September 2025Live
Most benchmarks ask a model to answer a test question. GDPval asks something closer to what an employer would ask: can the model produce a deliverable — a legal brief, an engineering drawing, a nursing care plan, a customer-support transcript — that a professional in that field would accept as good work? OpenAI built the tasks from 44 occupations across the nine sectors that contribute most to US GDP, each one drafted and checked by an industry professional with an average of 14 years’ experience, so the test items resemble real work products rather than academic exam questions.
Grading is blind: human experts compare a model’s output against an actual human-produced answer to the same task, without being told which is which. At launch, OpenAI reported that GPT-5 and Claude Opus 4.1 — the top two systems it tested — produced work rated as equal to or better than the human comparison on close to half of tasks, and that performance had been improving steadily across model generations. Only a 220-task “gold” subset was released publicly, graded through an automated evaluation service OpenAI hosts, rather than the full 1,320-task set.
GDPval arrived as part of a broader shift, visible the same month in Scale AI’s SWE-bench Pro, away from saturating academic tests and toward harder evaluations grounded in professional output. Its explicit framing around GDP-weighted occupations rather than a research subfield made it a reference point in debates about AI’s effect on knowledge work specifically, and third-party trackers such as Artificial Analysis have since built leaderboards on top of the public gold subset as newer models are released.
The set
1,320 tasks across 44 occupations drawn from the nine US industry sectors that contribute most to GDP, each built and vetted by a professional averaging 14 years' experience; a 220-task 'gold' subset is public, graded via an automated evaluation service OpenAI hosts. Outputs are compared blind against human-produced work on the same task.
Example
An Accountants and Auditors task from the public gold subset: 'You are an auditor and as part of an audit engagement, you are tasked with reviewing and testing the accuracy of reported Anti-Financial Crime Risk Metrics. The attached spreadsheet titled 'Population' contains Anti-Financial Crime Risk Metrics for Q2 and Q3 2024. ... Calculate the required sample size for audit testing based on a 90% confidence level and a 10% tolerable error rate...', supplied with a reference spreadsheet and graded against a human auditor's own deliverable.huggingface.co
Where it stands
Introduced in September 2025; a third-party GDPval-AA leaderboard (Artificial Analysis) now tracks new frontier models against the public gold subset.
In the timeline · 4 entries
ByteDance releases Seed 2.1 model family
The closed-weight Pro and Turbo models target multi-step agent tasks such as mobile-app control and coding, sold through Volcano Engine at prices quoted in yuan.
Models & capabilities
UC Berkeley releases Agents' Last Exam, a benchmark of professional work
Built with 250+ industry experts across 55 sub-industries, it runs agents in the real software a specialist would use and grades against hidden answers; current systems clear under 1% of the hardest tier.
Benchmarks & progress
OpenAI releases GPT-5.2
Released three weeks after Google's Gemini 3 and following a reported internal OpenAI 'code red,' with a claimed 70.9% win rate against professionals on the GDPval benchmark, up from 38.8% for GPT-5.1.
Models & capabilities
OpenAI publishes GDPval, a benchmark for economically valuable knowledge work
Blind grading by industry professionals rated GPT-5 and Claude Opus 4.1 outputs as equal to or better than human work on nearly half of the 1,320 tasks.
Benchmarks & progress