Benchmarks · Multimodal

SpatialViz

also: SpatialViz-Bench, Spatial Visualization Benchmark

Spatial visualization: whether a multimodal model can mentally imagine and manipulate visual structures that are not directly observable, across mental rotation, mental folding, visual penetration and mental animation.

Institute of Automation, Chinese Academy of Sciences (CASIA), with UCAS, SJTU, ShanghaiTech, Huawei Noah's Ark & UCLReleased 10 July 2025Live

SpatialViz-Bench isolates a specific human ability that image recognition does not require: mentally picturing a shape and transforming it in the mind’s eye. Built by researchers at the Chinese Academy of Sciences’ Institute of Automation with several university and industry collaborators, it splits spatial visualization into four sub-abilities — mental rotation, mental folding, visual penetration and mental animation — and generates 1,180 image questions across twelve tasks designed to test each.

The questions are produced programmatically rather than hand-collected, a deliberate choice: procedurally generated items are far less likely to have appeared in a model’s training data, so a high score is harder to attribute to memorisation. A typical task shows a stack of cubes and four candidate rotations and asks which one could not have been produced by rotating the original — a question a person answers by turning the shape over mentally, and one that a model cannot shortcut by pattern-matching against a caption.

SpatialViz is one of the vision evaluations aggregated on the Center for AI Safety’s AI Dashboard, which runs it as an independent third-party test. There the trajectory mirrors the other spatial benchmarks — from around 30% for GPT-4o, near chance on the harder sub-abilities, up to roughly 80% for GPT-5.6 Sol by mid-2026 — but the leading score still leaves the hardest mental-transformation tasks unsolved, and the authors frame the benchmark as a diagnostic of which specific spatial skills a model lacks rather than a single pass/fail number.

The set

1,180 programmatically generated image questions across 12 tasks (three per sub-ability), each with two or three difficulty levels. Questions are multiple-choice — usually with image answer options — and graded by accuracy against a ground-truth key. Generating the items procedurally is meant to keep them out of training data.

Example

A 3D-rotation task: shown an original stack of equal-sized cubes on the left and four candidate stacks on the right, the model must answer which option cannot be obtained by rotating the original — where two options are reachable by rotation, one is a mirror or altered stack, and the correct answer is the stack produced by removing a cube rather than rotating it.arxiv.org

Where it stands

Tracked by the independent CAIS AI Dashboard as a standardized third-party evaluation. Scores have climbed from the low 30s in 2024 to around 80% by mid-2026, led by OpenAI's GPT-5.6 Sol, without saturating.

How the top score changed hands

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. May 2024GPT-4o31.1%An early multimodal baseline on the CAIS AI Dashboard, close to chance on the harder sub-abilities.
  2. December 2024o347.7%
  3. August 2025GPT-554.1%
  4. November 2025Gemini 3 Pro63.2%
  5. April 2026GPT-5.574.2%
  6. June 2026GPT-5.6 Sol80.7%The highest SpatialViz score recorded on the CAIS dashboard here.

Current best: GPT-5.6 Sol — 80.7% On the independent CAIS AI Dashboard, which runs SpatialViz as a standardized third-party evaluation across frontier models.

More multimodal benchmarks