MathVista
Can a model do mathematical reasoning when the problem is given as a picture — a plot, a geometry diagram, a puzzle figure — rather than as text?
UCLA, University of Washington & Microsoft Research (Lu, Bansal, Galley, Gao et al.)Released 3 October 2023
Word problems are easy to make harder by taking away the words. MathVista, built by researchers at UCLA, the University of Washington and Microsoft Research, tests mathematical reasoning that depends on reading a picture first — a plotted function, a geometry diagram, a bar chart, a logic-puzzle figure — and only then doing the arithmetic or algebra the picture implies. The benchmark pools 6,141 examples from 28 existing maths-and-vision datasets plus three tasks the authors built themselves to cover gaps: puzzle figures, function plots and academic-paper diagrams.
At launch in October 2023, the gap between models and people was stark. GPT-4V, the best of twelve systems tested, scored 49.9%, comfortably ahead of Google’s Bard but still 10.4 points short of the 60.3% human baseline the authors measured. That gap closed fast, and then reasoning models blew past it: Claude 3.5 Sonnet reported 67.7% in June 2024, already above the original human baseline, and by the end of that year OpenAI’s o3 reported 86.8% on the “testmini” subset with high reasoning effort — more than 25 points clear of the human baseline. Alibaba’s Qwen2.5-VL-72B reported 74.8% in February 2025, the strongest open-weight figure but still short of o3.
Unlike some contemporaries, MathVista has not been superseded by a discrete “Pro” successor, but it has faded from frontier labs’ own headline evaluation reports; Google’s Gemini 3 Pro evaluation document from November 2025, for instance, reports a newer slate of multimodal and reasoning benchmarks and does not include a MathVista figure. With the top scores now well above the human baseline, the benchmark is treated as effectively saturated.
The set
6,141 examples combining 28 existing maths-related multimodal datasets with three tasks the authors built specifically for the benchmark: IQTest (logical reasoning over puzzle figures), FunctionQA (algebraic reasoning over plotted functions) and PaperQA (reasoning over figures from academic papers). Answers are multiple-choice or free-form numeric/text, checked by exact or normalised match; a 5,000-example 'testmini' subset is the one most models report.
Example
Is the function (f: R to R) injective? Choices: (A) Yes (B) No (Correct output: (B) No)arxiv.org
Where it stands
GPT-4V scored 49.9% at launch against a 60.3% human baseline; reasoning models cleared it well before the end of 2024, with OpenAI's o3 reaching 86.8% on testmini (high reasoning). Qwen2.5-VL-72B's 74.8% remained the strongest open-weight figure. Most 2025-26 frontier labs have stopped headlining MathVista, treating it as saturated.
How the top score changed hands
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- October 2023GPT-4V49.9% (testmini)Best of 12 models evaluated at launch, 15.1 points ahead of Google's Bard, but still 10.4 points short of the 60.3% human baseline.
- June 2024Claude 3.5 Sonnet67.7% (testmini)From Anthropic's own model card addendum, ahead of GPT-4o (63.8%) and Gemini 1.5 Pro (63.9%) on the same table.
- December 2024o171.0% (testmini)OpenAI's first reasoning model, past the human baseline.
- December 2024o386.8% (testmini)High reasoning; the highest independently reported testmini score recorded here.
- February 2025Qwen2.5-VL-72B74.8% (testmini)The strongest open-weight result, above GPT-4o (63.8%) and Claude 3.5 Sonnet (67.7%) but below o3.
Current best: o3 — 86.8% (testmini) MathVista testmini, high reasoning (OpenAI's figure was updated on 16 April 2025 after a system-prompt change). Well above the 60.3% human baseline; the top open-weight result is Qwen2.5-VL-72B at 74.8%.
In the timeline · 2 entries
Moonshot AI releases Kimi K1.5 reasoning model
Moonshot said its RL-trained model matched OpenAI's o1 on multimodal reasoning without Monte Carlo tree search, but it launched the same week as DeepSeek-R1 and drew far less attention.
Models & capabilities · Benchmarks & progress
xAI releases Grok-2
The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
Models & capabilities