Benchmarks · Agents, tools & computer use
OSWorld
also: OSWorld-Verified, OSWorld 2.0
Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?
University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of WaterlooReleased 11 April 2024Live
OSWorld, released in April 2024 by researchers at the University of Hong Kong, Salesforce, Carnegie Mellon and Waterloo, tests something narrower and harder than most web-agent benchmarks: can a model operate an entire computer? Its 369 tasks run inside real Ubuntu, Windows and macOS environments with genuine applications — spreadsheets, file managers, browsers, settings panels — and are graded by checking the machine’s actual end-state afterwards, so an agent has to see the screen, decide what to click, and get the real result right.
At launch the best model managed only 12% against a 72% human baseline, and computer-use agents remained a research curiosity through most of 2024. That changed fast once labs began shipping dedicated computer-use products: OpenAI’s Operator reached 38% in January 2025, Claude Sonnet 4.5 hit 61% by September, and through early 2026 successive releases pushed into the 65–75% range on the refreshed “OSWorld-Verified” task set — with OpenAI’s GPT-5.4 the first to claim it had passed the human baseline outright. By mid-2026 the top figures reached ~85%: Anthropic reported Claude Fable 5 and Mythos 5 at 85.0% in June, and later releases did not surpass it — Moonshot’s Kimi K3 reached 84.8% and Google’s Gemini 3.6 Flash 83.0%. Every one of these is a vendor’s own figure, since OSWorld-Verified has no single independent leaderboard.
That pace of saturation prompted the same research group to release OSWorld 2.0 in June 2026: 108 much longer tasks, many taking a skilled human well over an hour, covering seven professional domains rather than isolated GUI actions. Early results reset the difficulty curve — even the strongest tested model cleared only about a fifth of tasks — reprising the pattern seen across the field’s other benchmarks, where a measure saturates and is promptly replaced by a harder one built the same way.
The set
369 tasks (361 excluding ones with an external Google Drive dependency) spanning real desktop and web applications across Ubuntu, Windows and macOS, graded by execution-based checker functions against the machine's actual end-state. A July 2025 refresh, OSWorld-Verified, fixed community-reported errors in the task set and added faster AWS-based parallel evaluation; the version most releases now cite.
Example
Rename "Sheet 1" to "LARS Resources". Then make a copy of it. Place the copy before "Sheet 2" and rename it by appending a suffix "(Backup)"arxiv.org
Where it stands
Original OSWorld tasks moved from a 12% ceiling in 2024 to the top OSWorld-Verified figures reaching ~85% by mid-2026 (Claude Fable 5 / Mythos 5), well past the ~72% human baseline — prompting the same group to release OSWorld 2.0 in June 2026 with 108 much longer tasks (median 1.6 hours of human time), on which even the strongest models clear only around 20%. Every headline figure is a vendor self-report; there is no single independent OSWorld-Verified leaderboard.
Editions, and how each was led
OSWorld-2.0current
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- June 2026Claude Fable 536.1% (strict)An early point on the harder 2.0 set; Anthropic-reported, strict scoring.
- July 2026Claude Opus 539.6% (strict)Strict scoring, as listed in Anthropic's Fable 5.1 comparison table.
- September 2026Claude Fable 5.141.7% (strict)Strict scoring; 77.9% with partial credit.
- September 2026Claude Opus 5.548.7% (strict)The current top under strict scoring; Anthropic-reported in the system card.
Current best: Claude Opus 5.5 — 48.7% (strict) Anthropic's Opus 5.5 system card, strict scoring (81.8% with partial credit). The same table lists Fable 5.1 at 42.8% and Opus 5 at 37.2% strict.
OSWorld-2.0 (partial)
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- June 2026Claude Fable 572.9% (partial)An early point on the 2.0 set, partial-credit; Anthropic-reported.
- September 2026Claude Fable 5.177.9% (partial)Partial-credit scoring; Anthropic-reported.
- September 2026GPT-6 Astra72.6% (partial)OpenAI's launch table (offline set, v2026.08.08), partial score.
- September 2026Claude Opus 5.581.8% (partial)The current top under partial-credit scoring; Anthropic-reported.
Current best: Claude Opus 5.5 — 81.8% (partial) Anthropic's Opus 5.5 launch table, partial-credit scoring (48.7% strict). Earlier tables: Fable 5.1 77.9%, Opus 5 75.4%, GPT-6 Astra 72.6%, GPT-5.6 Sol 65.7%, Gemini 3.8 Flash 59.0%.
OSWorld-VerifiedSaturated
Best reported score at each point the lead changed hands. Tap a marker for the entry.
- April 2024Best model at release (original paper)12.24%Against a 72.36% human baseline on the same tasks.
- January 2025OpenAI Operator (CUA)38.1%Against 22.0% for Anthropic's Computer Use and a 72.4% human baseline.
- September 2025Claude Sonnet 4.561.4%
- February 2026GPT-5.3-Codex64.7%On OSWorld-Verified, roughly double GPT-5.2-Codex's 38.2%.
- February 2026Claude models (Anthropic's own figure)72.5%Cited by Anthropic when announcing its acquisition of Vercept, up from under 15% in late 2024.
- March 2026GPT-5.475%On OSWorld-Verified; OpenAI framed this as beating its own ~72.4% human baseline figure.
- June 2026Claude Fable 5 / Mythos 585.0%OSWorld-Verified, Anthropic-reported — the top figure on this edition. Kimi K3 reached 84.8% and Gemini 3.6 Flash 83.0% a month or more afterwards.
Current best: Claude Fable 5 / Mythos 5 — 85.0% OSWorld-Verified, Anthropic-reported (Fable 5 and Mythos 5 both 85.0%). Kimi K3 (84.8%, Moonshot) and Claude Opus 4.8 (83.4%) sit just behind, all above Gemini 3.6 Flash's 83.0% — against a ~72% human baseline. All are vendor self-reports on the same task set.
In the timeline · 14 entries
Anthropic releases Claude Opus 5.5
Anthropic said Opus 5.5 matches Fable 5.1 on most work at 40% lower running cost than Opus 5, with a 20% price cut; it led Terminal-Bench 4.0 and Artificial Analysis's GDPval-AA board.
Models & capabilities · Benchmarks & progress
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1
Anthropic's launch table put Fable 5.1 ahead of its own Fable 5 and Opus 5 and OpenAI's GPT-5.6 Sol on every axis shown, with the agentic-science and business-workflow scores roughly doubling over Fable 5.
Models & capabilities · Benchmarks & progress
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber
Google cut per-token cost and lifted coding and computer-use benchmark scores on its efficiency tier, while Gemini 3.5 Pro remained unreleased and still in partner testing.
Models & capabilities
Anthropic launches Claude Sonnet 5
Priced at $3/$15 per million input/output tokens against Opus 4.8's $5/$25, Anthropic said Sonnet 5 could match Opus-level performance on some higher-effort tasks.
Models & capabilities
OpenAI releases GPT-5.4
OpenAI's first general-purpose model with built-in computer-use, reported scoring 75% on OSWorld-Verified against 47.3% for GPT-5.2 and roughly 72% for human testers.
Models & capabilities
Anthropic acquires Vercept
Terms were undisclosed; Vercept will wind down its own product, and Anthropic cited Claude's OSWorld computer-use score rising from under 15% in late 2024 to 72.5%.
Money & business · Labs & people
Anthropic releases Claude Sonnet 4.6
Early testers preferred it to Sonnet 4.5 on coding tasks about 70% of the time, and to the larger Opus 4.5 about 59% of the time, at unchanged Sonnet pricing.
Models & capabilities
OpenAI releases GPT-5.3-Codex
OpenAI reported the model roughly doubled its predecessor's OSWorld-Verified computer-use score, from 38.2% to 64.7%, and was the first Codex model rated 'High capability' for cybersecurity tasks.
Models & capabilities
Anthropic ships Claude Sonnet 4.5
Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
Models & capabilities
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
OpenAI launches Operator
Built on a new Computer-Using Agent model layered on GPT-4o, it scored 38.1% on OSWorld against a 72.4% human baseline, and launched to $200-a-month Pro subscribers only.
Models & capabilities
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities
OSWorld benchmarks AI agents on real desktop computer tasks
369 tasks across real Ubuntu, Windows and macOS applications, graded on the machine's actual end-state; the best model at release solved 12% against a 72% human baseline.
Benchmarks & progress