Benchmarks · Agents, tools & computer use

OSWorld

also: OSWorld-Verified, OSWorld 2.0

Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?

University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of WaterlooReleased 11 April 2024Live

OSWorld, released in April 2024 by researchers at the University of Hong Kong, Salesforce, Carnegie Mellon and Waterloo, tests something narrower and harder than most web-agent benchmarks: can a model operate an entire computer? Its 369 tasks run inside real Ubuntu, Windows and macOS environments with genuine applications — spreadsheets, file managers, browsers, settings panels — and are graded by checking the machine’s actual end-state afterwards, so an agent has to see the screen, decide what to click, and get the real result right.

At launch the best model managed only 12% against a 72% human baseline, and computer-use agents remained a research curiosity through most of 2024. That changed fast once labs began shipping dedicated computer-use products: OpenAI’s Operator reached 38% in January 2025, Claude Sonnet 4.5 hit 61% by September, and through early 2026 successive releases pushed into the 65–75% range on the refreshed “OSWorld-Verified” task set — with OpenAI’s GPT-5.4 the first to claim it had passed the human baseline outright. By mid-2026 the top figures reached ~85%: Anthropic reported Claude Fable 5 and Mythos 5 at 85.0% in June, and later releases did not surpass it — Moonshot’s Kimi K3 reached 84.8% and Google’s Gemini 3.6 Flash 83.0%. Every one of these is a vendor’s own figure, since OSWorld-Verified has no single independent leaderboard.

That pace of saturation prompted the same research group to release OSWorld 2.0 in June 2026: 108 much longer tasks, many taking a skilled human well over an hour, covering seven professional domains rather than isolated GUI actions. Early results reset the difficulty curve — even the strongest tested model cleared only about a fifth of tasks — reprising the pattern seen across the field’s other benchmarks, where a measure saturates and is promptly replaced by a harder one built the same way.

The set

369 tasks (361 excluding ones with an external Google Drive dependency) spanning real desktop and web applications across Ubuntu, Windows and macOS, graded by execution-based checker functions against the machine's actual end-state. A July 2025 refresh, OSWorld-Verified, fixed community-reported errors in the task set and added faster AWS-based parallel evaluation; the version most releases now cite.

Example

Rename "Sheet 1" to "LARS Resources". Then make a copy of it. Place the copy before "Sheet 2" and rename it by appending a suffix "(Backup)"arxiv.org

Where it stands

Original OSWorld tasks moved from a 12% ceiling in 2024 to the top OSWorld-Verified figures reaching ~85% by mid-2026 (Claude Fable 5 / Mythos 5), well past the ~72% human baseline — prompting the same group to release OSWorld 2.0 in June 2026 with 108 much longer tasks (median 1.6 hours of human time), on which even the strongest models clear only around 20%. Every headline figure is a vendor self-report; there is no single independent OSWorld-Verified leaderboard.

Editions, and how each was led

OSWorld-2.0current

Released June 2026108 much longer tasks (median ~1.6 hours of human work) across seven professional domains, released June 2026 after Verified saturated. Under strict scoring even the strongest models clear only around 40%. (Vendors also report a partial-credit number roughly 35 points higher — a different metric, kept off this axis.)

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Fable 536.1% (strict)An early point on the harder 2.0 set; Anthropic-reported, strict scoring.
  2. July 2026Claude Opus 539.6% (strict)Strict scoring, as listed in Anthropic's Fable 5.1 comparison table.
  3. September 2026Claude Fable 5.141.7% (strict)Strict scoring; 77.9% with partial credit.
  4. September 2026Claude Opus 5.548.7% (strict)The current top under strict scoring; Anthropic-reported in the system card.

Current best: Claude Opus 5.5 — 48.7% (strict) Anthropic's Opus 5.5 system card, strict scoring (81.8% with partial credit). The same table lists Fable 5.1 at 42.8% and Opus 5 at 37.2% strict.

OSWorld-2.0 (partial)

Released June 2026The same 2.0 task set scored with partial credit rather than strict pass/fail — the number vendors more often headline, and now reported across OpenAI, Anthropic and Google, so it earns its own axis. Runs roughly 30–35 points above the strict score for the same model.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Fable 572.9% (partial)An early point on the 2.0 set, partial-credit; Anthropic-reported.
  2. September 2026Claude Fable 5.177.9% (partial)Partial-credit scoring; Anthropic-reported.
  3. September 2026GPT-6 Astra72.6% (partial)OpenAI's launch table (offline set, v2026.08.08), partial score.
  4. September 2026Claude Opus 5.581.8% (partial)The current top under partial-credit scoring; Anthropic-reported.

Current best: Claude Opus 5.5 — 81.8% (partial) Anthropic's Opus 5.5 launch table, partial-credit scoring (48.7% strict). Earlier tables: Fable 5.1 77.9%, Opus 5 75.4%, GPT-6 Astra 72.6%, GPT-5.6 Sol 65.7%, Gemini 3.8 Flash 59.0%.

OSWorld-VerifiedSaturated

Released July 2025The 2024 original, refreshed in July 2025 to fix task errors and add faster parallel evaluation — the set most 2025–26 releases cite. Saturated past the ~72% human baseline by mid-2026, which prompted OSWorld 2.0.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. April 2024Best model at release (original paper)12.24%Against a 72.36% human baseline on the same tasks.
  2. January 2025OpenAI Operator (CUA)38.1%Against 22.0% for Anthropic's Computer Use and a 72.4% human baseline.
  3. September 2025Claude Sonnet 4.561.4%
  4. February 2026GPT-5.3-Codex64.7%On OSWorld-Verified, roughly double GPT-5.2-Codex's 38.2%.
  5. February 2026Claude models (Anthropic's own figure)72.5%Cited by Anthropic when announcing its acquisition of Vercept, up from under 15% in late 2024.
  6. March 2026GPT-5.475%On OSWorld-Verified; OpenAI framed this as beating its own ~72.4% human baseline figure.
  7. June 2026Claude Fable 5 / Mythos 585.0%OSWorld-Verified, Anthropic-reported — the top figure on this edition. Kimi K3 reached 84.8% and Gemini 3.6 Flash 83.0% a month or more afterwards.

Current best: Claude Fable 5 / Mythos 5 — 85.0% OSWorld-Verified, Anthropic-reported (Fable 5 and Mythos 5 both 85.0%). Kimi K3 (84.8%, Moonshot) and Claude Opus 4.8 (83.4%) sit just behind, all above Gemini 3.6 Flash's 83.0% — against a ~72% human baseline. All are vendor self-reports on the same task set.

In the timeline · 14 entries

More agents, tools & computer use benchmarks