ARC Prize Foundation launches ARC-AGI-3
Humans scored 100% and frontier AI scored 0.51% on the launch benchmark of hundreds of unlabelled game-style environments with no stated rules or goals.
- Benchmarks & progress
- Notable
The ARC Prize Foundation launched ARC-AGI-3, the third generation of its ARC-AGI benchmark, moving away from the static, one-shot visual puzzles of the previous two versions toward hundreds of interactive, turn-based game environments handcrafted by human designers. Agents are given no instructions, rules or stated goals; they have to explore an unfamiliar environment, work out its mechanics, and demonstrate they can act on what they learned, rather than pattern-match against training data. The environments are playable in-browser or via API and form the core of ARC Prize 2026, a competition offering over $2 million in prizes.
The foundation reported that humans scored 100% on the launch set, against 0.51% for frontier AI systems — a gap the organisers presented as evidence that current models remain poor at open-ended exploration and rapid skill acquisition, in contrast to the narrower gap that had opened up on ARC-AGI-2’s static puzzles. The shift followed a familiar pattern in the ARC series: each generation targeted the kind of task that models had begun to do well on, moving the goalposts to isolate a capability — here, learning a new game’s rules from interaction alone — that pattern-matching on static problems does not require.
The launch event, hosted by Y Combinator in San Francisco, paired ARC-AGI creator François Chollet with OpenAI’s Sam Altman in a discussion of how to measure progress toward artificial general intelligence, underscoring how central the benchmark had become to the industry’s own account of what remained unsolved. A follow-up study establishing human performance baselines across the full set of environments and revising the scoring methodology followed in April.