Timeline

OpenAI publishes GDPval, a benchmark for economically valuable knowledge work

Blind grading by industry professionals rated GPT-5 and Claude Opus 4.1 outputs as equal to or better than human work on nearly half of the 1,320 tasks.

  • Benchmarks & progress
  • Notable

OpenAI published GDPval, a benchmark intended to measure model performance against real, economically valuable knowledge work rather than academic test questions. The set comprised 1,320 tasks across 44 occupations chosen from industries each contributing more than 5% of US GDP, with a 220-task “gold” subset released openly. Each task — a legal brief, an engineering blueprint, a customer support transcript, a nursing care plan — was built and vetted by professionals with an average of 14 years’ experience in the relevant field, and modelled on an actual work product rather than a synthetic question.

Outputs were graded blind by industry experts comparing model responses against human-produced work on the same task, without knowing which was which. OpenAI reported that its GPT-5 and Anthropic’s Claude Opus 4.1 — the two top-performing systems tested — produced work rated as equal to or better than the human comparison in close to half of tasks.

The benchmark arrived amid a wider push, visible also in Scale AI’s SWE-bench Pro released the same month, to replace saturating academic benchmarks with harder, more realistic evaluations grounded in actual professional output. Its framing was explicitly economic rather than purely technical: by anchoring task selection to GDP share rather than to a research subfield, OpenAI positioned GDPval as evidence for the argument that model capability was approaching a threshold relevant to the labour market, not just to leaderboard rankings — a claim that in turn fed directly into ongoing debate about AI’s effect on knowledge-worker employment.

Referenced by