MindCube
also: MindCube benchmark, Spatial Mental Modeling
Whether a vision-language model can build a spatial mental model of a scene — inferring the positions, orientations and possible movements of objects, including ones it cannot currently see — from only a few limited views.
Northwestern University (MLL-Lab), with Stanford, NYU & University of WashingtonReleased 26 June 2025Live
MindCube asks a question that is easy for people and hard for machines: shown a scene from only a couple of angles, can a model work out where everything is — including the things currently out of frame — and how the layout would look if it moved? Built by Manling Li’s lab at Northwestern with collaborators at Stanford, NYU and the University of Washington, the benchmark pairs 3,268 images into 976 multi-view sets and asks more than 21,000 multiple-choice questions across three settings — Rotation, Around and Among — each demanding a different mental transformation of the scene.
The design deliberately withholds the convenient view. A model cannot answer “what is to the left of the white jar, from the viewpoint in image 4?” by reading a single picture; it has to fuse the available views into an internal spatial model and then query that model from a perspective it was never shown. That is what the authors mean by a “spatial mental model”, and it is a capability distinct from recognising objects or captioning an image.
MindCube is one of the vision evaluations aggregated on the Center for AI Safety’s AI Dashboard, which runs it as a standardized third-party test. On that dashboard the climb has been steep — from the high 40s for GPT-4o to the mid-80s by mid-2026 — but the frontier is now closely bunched, with Claude Opus 5 and Gemini 3.1 Pro separated by two-tenths of a point. Whether the remaining gap to the top reflects genuinely harder spatial reasoning or the benchmark approaching its ceiling is not yet settled.
The set
21,154 multiple-choice questions over 3,268 still images grouped into 976 multi-view sets, each set being two to four viewpoints of a single scene (multiple images per question, not video). Three settings — Rotation, Around and Among — probe different mental-transformation demands; grading is question-answering accuracy.
Example
From the viewpoint presented in image 4, what is to the left of the white jar? A. Table with cups on it B. Clothes rack C. Bed sheet with a floral pattern D. White headboard — a Rotation-setting question in which the model sees a white jar from several viewpoints and must reason about a viewpoint it must reconstruct. (Answer: C.)arxiv.org
Where it stands
Tracked by the independent CAIS AI Dashboard as a standardized third-party evaluation. Scores have risen from the high 40s in 2024 to the mid-80s by mid-2026, with the frontier tightly bunched at the top.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- May 2024GPT-4o48.1%An early multimodal baseline on the CAIS AI Dashboard.
- December 2024o357.6%
- July 2025Grok 464.7%
- November 2025Gemini 3 Pro77.3%A large jump, past three-quarters.
- February 2026Gemini 3.1 Pro84.1%
- July 2026Claude Opus 584.3%The highest MindCube score recorded on the CAIS dashboard here.
Current best: Claude Opus 5 — 84.3% On the independent CAIS AI Dashboard, narrowly ahead of Gemini 3.1 Pro (84.1%) — the frontier is closely bunched here.