Benchmarks · Agents, tools & computer use

OSWorld

also: OSWorld-Verified, OSWorld 2.0

Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?

University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of WaterlooReleased 11 April 2024Live

OSWorld, released in April 2024 by researchers at the University of Hong Kong, Salesforce, Carnegie Mellon and Waterloo, tests something narrower and harder than most web-agent benchmarks: can a model operate an entire computer? Its 369 tasks run inside real Ubuntu, Windows and macOS environments with genuine applications — spreadsheets, file managers, browsers, settings panels — and are graded by checking the machine’s actual end-state afterwards, so an agent has to see the screen, decide what to click, and get the real result right.

At launch the best model managed only 12% against a 72% human baseline, and computer-use agents remained a research curiosity through most of 2024. That changed fast once labs began shipping dedicated computer-use products: OpenAI’s Operator reached 38% in January 2025, Claude Sonnet 4.5 hit 61% by September, and by early 2026 successive point releases from OpenAI and Google were reporting scores in the 65–83% range on the refreshed “OSWorld-Verified” task set — with OpenAI’s GPT-5.4 the first to claim it had passed the human baseline outright.

That pace of saturation prompted the same research group to release OSWorld 2.0 in June 2026: 108 much longer tasks, many taking a skilled human well over an hour, covering seven professional domains rather than isolated GUI actions. Early results reset the difficulty curve — even the strongest tested model cleared only about a fifth of tasks — reprising the pattern seen across the field’s other benchmarks, where a measure saturates and is promptly replaced by a harder one built the same way.

The set

369 tasks (361 excluding ones with an external Google Drive dependency) spanning real desktop and web applications across Ubuntu, Windows and macOS, graded by execution-based checker functions against the machine's actual end-state. A July 2025 refresh, OSWorld-Verified, fixed community-reported errors in the task set and added faster AWS-based parallel evaluation; the version most releases now cite.

Example

Rename "Sheet 1" to "LARS Resources". Then make a copy of it. Place the copy before "Sheet 2" and rename it by appending a suffix "(Backup)"arxiv.org

Where it stands

Original OSWorld tasks moved from a 12% ceiling in 2024 to frontier models matching or beating the ~72% human baseline by mid-2026, prompting the same group to release OSWorld 2.0 in June 2026 with 108 much longer tasks (median 1.6 hours of human time) — on which even the strongest models clear only around 20%.

How the top score changed hands

  1. April 2024Best model at release (original paper)12.24%Against a 72.36% human baseline on the same tasks.
  2. January 2025OpenAI Operator (CUA)38.1%Against 22.0% for Anthropic's Computer Use and a 72.4% human baseline.
  3. September 2025Claude Sonnet 4.561.4%
  4. February 2026GPT-5.3-Codex64.7%On OSWorld-Verified, roughly double GPT-5.2-Codex's 38.2%.
  5. February 2026Claude models (Anthropic's own figure)72.5%Cited by Anthropic when announcing its acquisition of Vercept, up from under 15% in late 2024.
  6. March 2026GPT-5.475%On OSWorld-Verified; OpenAI framed this as beating its own ~72.4% human baseline figure.
  7. July 2026Gemini 3.6 Flash83.0%

Current best: Gemini 3.6 Flash — 83.0% On OSWorld-Verified, as reported in Google's release announcement; against a roughly 72% human baseline on the same task set.

In the timeline · 11 entries

More agents, tools & computer use benchmarks