Timeline

ARC Prize publishes ARC-AGI-3 human performance dataset

The 458-participant study replaced a second-best-player baseline with the median player, reducing the effect of luck on any single level's score.

  • Benchmarks & progress
  • Minor

The ARC Prize Foundation published a controlled study establishing human performance baselines across all 135 environments in its ARC-AGI-3 benchmark, drawn from weekly in-person focus groups of 458 participants run out of a San Francisco testing centre — the foundation’s most extensive human study in the ARC-AGI series to date. It released the underlying data as a public dataset of 342 step-by-step human replays covering the 25 environments made publicly available.

The study introduced the series’ first formal measure of learning efficiency: rather than only recording whether a level was completed, it tracked how many actions a solver needed relative to the human median, with 100% indicating a solve at the median action count. That efficiency measure fed into two changes to how ARC-AGI-3 is scored. The comparison baseline shifted from the second-best human player to the median player on each level, reducing the degree to which a single lucky or unlucky run could swing a score, and the per-level score cap was raised from 100% to 115%, so that one badly performed level could no longer drag an overall score down disproportionately.

The foundation said the recalibration drew on nearly one million scorecards already submitted against the benchmark’s public environments, using that volume of real submissions to stress-test the new scoring approach before publishing it. The change gave researchers a firmer basis for comparing AI systems’ sample efficiency — not just whether a model could eventually solve a task, but how many attempts it needed relative to a human baseline.