Benchmarks · Multimodal

RealWorldQA

Does a model's visual understanding hold up on ordinary real-world photos — many taken from inside or around a vehicle — that test spatial reasoning and everyday physical common sense, rather than on curated benchmark imagery?

xAIReleased April 2024Live

RealWorldQA is unusual among multimodal benchmarks in how little apparatus surrounds it: xAI released it as a plain dataset alongside its Grok-1.5 Vision preview in April 2024, with no accompanying paper and no public leaderboard. Its questions come from ordinary real-world photographs — many taken from inside or around a vehicle — paired with a single verifiable answer, testing spatial reasoning and physical common sense rather than the curated, often text-heavy imagery that dominates academic multimodal benchmarks. The initial release held just over 700 images, and xAI said it intended to grow the set over time.

Because there is no formal leaderboard, RealWorldQA’s numbers travel the way informal benchmarks often do in this record: cited by whichever lab wants to make a comparative claim, in whatever table its own release material assembles. The clearest multi-model comparison available is a third-party one, in Alibaba’s Qwen2.5-VL technical report from February 2025, which put InternVL2.5-78B and the benchmark’s previous open-source state of the art tied at 78.7%, just ahead of Qwen2-VL-72B (77.8%) and GPT-4o (75.4%). Claude 3.5 Sonnet trailed noticeably on the same table, at 60.1% — a large enough gap to suggest the benchmark’s vehicle-and-scene imagery favours some training distributions over others.

RealWorldQA’s appeal is precisely its informality: it asks whether a model’s visual competence generalises to the kind of photograph nobody curated for a benchmark, which is also why it resists the kind of authoritative, single-source leaderboard that benchmarks with a dedicated maintaining team can offer. It continues to appear as a comparison point in vision-language model release material, but — unlike MMMU or ChartQA — with no canonical current leader that this entry could confirm.

The set

An initial release of more than 700 real-world photographs, many anonymised images taken from vehicles, each paired with a question and a single easily-verifiable answer. xAI released it as a plain dataset rather than a formal paper, and said at launch it intended to expand the set over time.

Example

Which of the 3 objects is the smallest? A. The object on the right is the smallest object. B. The object on the left is the smallest object. C. The object in the middle is the smallest object. (with an accompanying image; answer: C)huggingface.co

Where it stands

Released without a leaderboard or formal paper; scores are self-reported by whichever lab cites the benchmark. By February 2025, InternVL2.5-78B and the original open-source state of the art were tied for the highest score in one widely-cited third-party comparison table, with Claude 3.5 Sonnet notably behind the rest of the field.

How the top score changed hands

  1. February 2025Claude 3.5 Sonnet / GPT-4o / InternVL2.5-78B / Qwen2-VL-72B / Qwen2.5-VL-72B60.1% / 75.4% / 78.7% / 77.8% / 75.7%From Alibaba's Qwen2.5-VL technical report — the only precisely-dated, multi-model comparison found for this entry, since xAI released the benchmark without an accompanying paper or public leaderboard.

Current best: InternVL2.5-78B — 78.7% Tied with the 'previous open-source SoTA' figure in the same comparison table, ahead of Qwen2-VL-72B (77.8%), Qwen2.5-VL-72B (75.7%) and GPT-4o (75.4%); Claude 3.5 Sonnet trailed at 60.1% on the same table. Not from an official RealWorldQA leaderboard, since none exists; later 2025-26 frontier models plausibly score higher but are not confirmed here.

More multimodal benchmarks