Benchmarks · Multimodal

IntPhys 2

also: IntPhys2, Intuitive Physics 2

Intuitive physics from video: whether a model grasps four macroscopic principles — object permanence, immutability, spatio-temporal continuity and solidity — well enough to tell physically possible events from impossible ones.

FAIR at MetaReleased 11 June 2025Live

IntPhys 2 tests whether a model has the intuitive grasp of physics that human infants develop early: that objects persist when hidden, keep their identity, move continuously through space, and cannot pass through solid matter. Built by FAIR at Meta as the successor to the 2019 IntPhys, it uses the “violation-of-expectation” method borrowed from developmental psychology — pairing an ordinary scene with a physically impossible version of the same scene, and asking whether a system notices the difference.

The task is set up so fluency cannot fake it. A model watches short synthetic videos and classifies each as physically possible or impossible; because the possible and impossible versions are matched, a model that has genuinely modelled the scene’s physics should split them, while one relying on surface cues should not. The main split holds 253 scenes across Easy, Medium and Hard difficulty, and the headline result at release was a wide gap: human testers score about 96%, the best models around 57%, and many barely clear the 50% chance line.

IntPhys 2 is one of the evaluations aggregated on the Center for AI Safety’s AI Dashboard, which runs it independently across frontier models. Unlike the spatial-reasoning benchmarks alongside it, IntPhys 2 has moved slowly: the top score crept from the low 50s to the mid-60s over two years, with Claude Fable 5 leading at 67.1% by mid-2026 — still some thirty points short of people. That persistent gap is the benchmark’s argument: models that can describe a physical scene in fluent prose still do not reliably know what is physically possible within it.

The set

A video benchmark built on the violation-of-expectation paradigm: scenes are arranged in matched sets of possible and impossible versions, and a multimodal model must classify each as plausible or implausible. The main evaluation split holds 253 synthetic scenes (1,012 videos), sub-divided into Easy, Medium and Hard. Human testers reach about 96%; the strongest models sit near 57–67% and many hover close to the 50% chance line.

Example

A solidity test: a ball rolls down a slope toward a brick. In the possible video the ball strikes the brick and changes course; in the impossible counterpart it passes straight through the solid brick without altering its path. The model must flag the pass-through version as physically impossible.arxiv.org

Where it stands

Tracked by the independent CAIS AI Dashboard as a standardized third-party evaluation. Progress is slow and the frontier remains far below the ~96% human baseline — top models reach only the mid-60s, with many near the 50% chance line.

How the top score changed hands

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. May 2024GPT-4o53%Barely above the 50% chance line on the CAIS AI Dashboard.
  2. March 2025Gemini 2.5 Pro56%
  3. December 2025Gemini 3 Flash63.4%The first clearly past 60%.
  4. May 2026Gemini 3.5 Flash67%
  5. June 2026Claude Fable 567.1%The highest IntPhys 2 score recorded on the CAIS dashboard here, still far short of human performance.

Current best: Claude Fable 5 — 67.1% On the independent CAIS AI Dashboard, narrowly ahead of Gemini 3.5 Flash (67.0%) — still roughly 30 points below the human baseline of about 96%.

More multimodal benchmarks