Timeline

Benchmark tests image and video generation as a form of reasoning

Using rule-based rather than AI-judge scoring across 300 tasks, the authors found video generation strongest at spatiotemporal tracking and interleaved generation the most compute-efficient.

  • Benchmarks & progress
  • Models & capabilities
  • Minor

A 52-author team spanning many institutions — including Dahua Lin of the Chinese University of Hong Kong and Zhongang Cai of SenseTime Research — published VBVR-Pro, a testbed for what the authors call “native visual reasoning”: using image and video generation itself, rather than text, as the medium in which a model works through a problem, treating generated visual states as steps in reasoning rather than only as a finished output to be judged.

The benchmark comprises 300 procedurally generated tasks, scored with deterministic, task-specific rules rather than by having another vision-language model judge the output — a choice the authors made because model-based judging is itself unreliable and can introduce its own bias into a benchmark meant to measure reasoning. Running more than 30 generation systems through controlled comparisons, the authors reported that video-generation approaches performed strongest on tasks requiring tracking of objects through time and space, while interleaved generation — alternating between generated images and text within a single reasoning process — was the more compute-efficient approach among those tested. The authors also checked whether performance on their tasks transferred to seven existing external benchmarks, and reported that it did.

The paper adds a rule-scored, generation-centred alternative to the more common practice of testing multimodal models only on their ability to describe or answer questions about images and video shown to them, at a moment when video-generation systems — such as DeepMind’s Genie 3 world model — are increasingly pitched as tools for simulating and planning within a world rather than only for producing finished media.