ARC Prize 2025 results and analysis published
The Kaggle track's top score reached 24% on ARC-AGI-2 within the competition's cost limits, while Gemini 3 Pro scored around 54% unconstrained, using iterative test-time refinement.
- Benchmarks & progress
- Minor
The ARC Prize Foundation published results and analysis for the 2025 edition of its ARC-AGI reasoning competition, run on ARC-AGI-2, the harder benchmark it introduced the previous year.
The Kaggle track, which caps compute cost per task, drew 1,455 teams and 15,154 submissions, similar participation to 2024. The winning submission, from team NVARC, reached 24.03% accuracy at an average cost of $0.20 per task; the ARChitects and MindsAI placed second and third. Evaluated separately, without the competition’s cost constraints, frontier chat models scored far higher — Gemini 3 Pro reached around 54% using an application-layer refinement approach — illustrating a gap between compute-efficient, purpose-built solutions and unconstrained use of large general models on the same benchmark. The foundation’s paper track drew 90 submissions, up from 47 the year before; the winning paper, a “Tiny Recursive Model” using only 7 million parameters, reached 45% on the original, easier ARC-AGI-1.
The organisers framed “refinement loops” — iterative exploration and verification of candidate solutions — as the dominant theme of 2025’s progress on the benchmark, and argued that reliable feedback signals and knowledge coverage, rather than raw model scale alone, were now the main limiting factors on further gains. The competition’s grand prize, requiring a higher accuracy threshold within its cost constraints, remained unclaimed for the second year running.