Timeline

ARC Prize compares reasoning models with no clear winner

ARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.

  • Benchmarks & progress
  • Minor

The ARC Prize Foundation published a comparison of frontier reasoning models across its ARC-AGI-1 and ARC-AGI-2 benchmarks, testing systems from OpenAI (o3, o3-pro, o4-mini), Anthropic (Claude Sonnet 4, Claude Opus 4), Google (Gemini 2.5 Flash and Pro), DeepSeek (R1), xAI (Grok 3) and Meta (Llama 4). Its conclusion was that no model was straightforwardly best: the right choice depended on whether a user prioritised raw accuracy, cost per task, or a balance of the two, with different models leading on each axis.

OpenAI’s higher-effort o3 configurations led on accuracy regardless of cost. Google’s Gemini 2.5 Flash offered the best balance of accuracy against price, and xAI’s cheaper Grok 3 mini configuration delivered modest gains over older, non-reasoning LLMs at low cost. None of the systems tested solved ARC-AGI-2, the harder of the two benchmarks, which the foundation had designed specifically to resist the kind of brute-force search and memorisation that had allowed earlier models to make rapid gains on ARC-AGI-1.

The report’s framing — that current reasoning systems shared common limitations regardless of vendor — argued against a simple “leaderboard” view of AI progress in which one model’s benchmark score settles the question of which is most capable. ARC Prize positioned the unsolved ARC-AGI-2 benchmark as evidence that scaling up existing reasoning approaches was not, on its own, sufficient to close the remaining gap toward more general problem-solving.