OpenAI introduces FrontierScience benchmark
GPT-5.2 scored 77% on olympiad-style questions but 25% on open-ended research tasks, a gap OpenAI's own researchers said showed little improvement over GPT-5.
- Benchmarks & progress
- Minor
OpenAI released FrontierScience, a benchmark of expert-written questions in physics, chemistry and biology, split into an olympiad-style tier of roughly 100 problems and a smaller, harder research tier of about 60 open-ended questions designed by PhD scientists. Lead researcher Miles Wang said the aim was to measure whether models could approach — and eventually accelerate — the work of practising scientists, rather than test recall of established facts.
OpenAI reported GPT-5.2 topping both tiers among frontier models, scoring 77.1% on the olympiad-style questions but only 25.3% on the research-tier questions, and said the improvement over the preceding GPT-5 was negligible on the research tier specifically. TIME reported that the benchmark provided no human baseline for comparison, covered text only — excluding experimental work or interpretation of images and video — and that its small question count limited how confidently models could be ranked against each other.
The gap between structured and open-ended performance was itself the notable result: it offered evidence, from OpenAI’s own evaluation, that benchmark gains on well-defined problems had not yet translated into comparable gains on the kind of ill-defined, exploratory reasoning that scientific research actually requires.