Timeline

ARC Prize analyses o3 and o4-mini on ARC-AGI

The publicly shipped o3 scored 41-53% on ARC-AGI-1, far below the 76-88% OpenAI's pre-release preview had shown the previous December.

  • Benchmarks & progress
  • Models & capabilities
  • Minor

The ARC Prize Foundation published an analysis comparing the ARC-AGI scores of OpenAI’s publicly released o3 and o4-mini models against the scores OpenAI itself had reported for a pre-release version of o3 the previous December. The gap was large: the December preview had scored 76% at low compute and 88% at high compute on ARC-AGI-1, while the production model that shipped in April scored 41% at low reasoning effort and 53% at medium — with both configurations under 3% on the harder ARC-AGI-2 set. Production o4-mini scored lower still, 21% and 41% at its two effort settings.

ARC Prize attributed the gap to several changes between the preview and the shipped product: the production o3 appeared to be a different underlying model, ran with less test-time compute than the December version had access to, gained multimodal input, and had been fine-tuned for chat and general product use rather than for the benchmark. The analysis also reported per-task inference cost — a few dollars per task at the higher reasoning setting — as a companion figure to the accuracy scores, underlining that headline reasoning-model results depend heavily on how much compute is spent per problem.

The write-up became a widely cited example of the distance between a lab’s pre-release capability claims and the scores its shipped product actually earns, feeding into broader scepticism about benchmark numbers announced ahead of a general release.