Benchmarks · Agents, tools & computer use

Online-Mind2Web

also: Mind2Web, Online Mind2Web

How well does a web agent actually perform on real, live websites under conditions close to genuine use — and how much of the field's reported progress on other web-agent benchmarks holds up under an independent, harder check?

Ohio State University NLP group, with UC BerkeleyReleased 2 April 2025Disputed

Mind2Web, released by an Ohio State University team in 2023, was one of the first large datasets built to train and test “generalist” web agents — over 2,000 tasks across 137 real websites and 31 domains, recorded as page snapshots rather than run live. Two years later, much of the same group returned with a pointed successor: Online-Mind2Web, introduced in a paper titled “An Illusion of Progress? Assessing the Current State of Web Agents,” built specifically to check whether the field’s rapid-looking gains on other benchmarks reflected real capability.

The answer, in the paper’s own testing, was: not entirely. Re-evaluating widely cited agents under careful, consistent human judgment rather than each lab’s own self-reported numbers, the authors found OpenAI’s Operator — then the field’s strongest browsing agent by most other measures — completing only 61.3% of tasks, with Claude’s Computer Use trailing at 56.3% and most other tested agents scoring around 30%, well below the figures those same systems had been credited with elsewhere.

Online-Mind2Web has since become a check against inflated web-agent claims rather than a benchmark labs race to top: Princeton’s independently run HAL leaderboard, which tracked submissions through 2025, reported the best agents clearing only around 40% and has since paused adding new models to focus on measuring reliability instead of chasing a single accuracy number. Its central argument — that a benchmark’s headline score depends heavily on who is doing the grading — has fed directly into wider scepticism about self-reported agent results across the field.

The set

300 realistic tasks across 136 live websites, evaluated with an LLM-as-judge method the authors report agrees with human judgment about 85% of the time; the paper also re-tests several widely cited agents under consistent, careful human evaluation rather than relying on each lab's self-reported numbers. It succeeds the original Mind2Web (2023), a static dataset of over 2,000 tasks across 137 websites and 31 domains built for training and evaluating web agents on recorded page snapshots rather than live sites.

Example

Compare Audi A7 with Audi A6 both made in 2023 and hide similaritiesarxiv.org

Where it stands

The paper introducing this benchmark, titled 'An Illusion of Progress? Assessing the Current State of Web Agents,' argues several widely cited web-agent results do not hold up under independent, careful evaluation; Princeton's HAL project, which ran an open leaderboard for it, has since paused adding new models to focus on measuring reliability rather than headline accuracy, and reports leading agents still only clearing around 40%.

How the top score changed hands

  1. April 2025OpenAI Operator61.3%Best-performing agent in the original paper's independent evaluation; Claude Computer Use 3.7 followed at 56.3%, and most other tested agents scored around 30%.
  2. August 2025SeeAct with GPT-5 (medium)42.33%Top entry on Princeton's independently run HAL leaderboard, well below labs' own self-reported figures on other web benchmarks.

Current best: OpenAI Operator — 61.3% The paper's own independent evaluation of Operator, presented as markedly lower than headline figures Operator had been given on other web-agent benchmarks.

In the timeline · 1 entry

More agents, tools & computer use benchmarks