Benchmarks · Multimodal

MathVista

Can a model do mathematical reasoning when the problem is given as a picture — a plot, a geometry diagram, a puzzle figure — rather than as text?

UCLA, University of Washington & Microsoft Research (Lu, Bansal, Galley, Gao et al.)Released 3 October 2023Saturating

Word problems are easy to make harder by taking away the words. MathVista, built by researchers at UCLA, the University of Washington and Microsoft Research, tests mathematical reasoning that depends on reading a picture first — a plotted function, a geometry diagram, a bar chart, a logic-puzzle figure — and only then doing the arithmetic or algebra the picture implies. The benchmark pools 6,141 examples from 28 existing maths-and-vision datasets plus three tasks the authors built themselves to cover gaps: puzzle figures, function plots and academic-paper diagrams.

At launch in October 2023, the gap between models and people was stark. GPT-4V, the best of twelve systems tested, scored 49.9%, comfortably ahead of Google’s Bard but still 10.4 points short of the 60.3% human baseline the authors measured. That gap closed fast, and then reasoning models blew past it: Claude 3.5 Sonnet reported 67.7% in June 2024, already above the original human baseline, and by the end of that year OpenAI’s o3 reported 86.8% on the “testmini” subset with high reasoning effort — more than 25 points clear of the human baseline. Alibaba’s Qwen2.5-VL-72B reported 74.8% in February 2025, the strongest open-weight figure but still short of o3.

Unlike some contemporaries, MathVista has not been superseded by a discrete “Pro” successor, but it has faded from frontier labs’ own headline evaluation reports; Google’s Gemini 3 Pro evaluation document from November 2025, for instance, reports a newer slate of multimodal and reasoning benchmarks and does not include a MathVista figure. With the top scores now well above the human baseline, the benchmark is treated as effectively saturated.

The set

6,141 examples combining 28 existing maths-related multimodal datasets with three tasks the authors built specifically for the benchmark: IQTest (logical reasoning over puzzle figures), FunctionQA (algebraic reasoning over plotted functions) and PaperQA (reasoning over figures from academic papers). Answers are multiple-choice or free-form numeric/text, checked by exact or normalised match; a 5,000-example 'testmini' subset is the one most models report.

Example

Is the function (f: R to R) injective? Choices: (A) Yes (B) No (Correct output: (B) No)arxiv.org

Where it stands

GPT-4V scored 49.9% at launch against a 60.3% human baseline; reasoning models cleared it well before the end of 2024, with OpenAI's o3 reaching 86.8% on testmini (high reasoning). Qwen2.5-VL-72B's 74.8% remained the strongest open-weight figure. Most 2025-26 frontier labs have stopped headlining MathVista, treating it as saturated.

How the top score changed hands

In the timeline · 2 entries

More multimodal benchmarks