OSWorld 2.0
Also known as OSWorld-V2, OSWorld2.0
OSWorld 2.0 tests whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state. It exists because agents had reached 83.5% on the original OSWorld, and it shows how much of that progress was benchmark-specific: the best full-completion score here is about 20%.
| What it measures | Whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state. |
|---|---|
| Built by | XLANG Lab, University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and others, 2026 |
| Format | 108 long-horizon tasks across 7 professional domains and 21 sub-categories, using 31 self-hosted websites alongside desktop applications and authentic input files |
| Scoring metric | Binary completion at a 500-step cap, reported with a weighted-checkpoint partial score |
| Status | Active |
| Representative top score | 20.6% · Claude Opus 4.8 (max thinking) · read 2026-06 |
| Official leaderboard | osworld-v2.xlang.ai |
How OSWorld 2.0 works
The benchmark contains 108 long-horizon tasks across seven professional domains and 21 subcategories, using 31 self-hosted websites alongside real desktop applications and authentic input files. Grading is binary completion at a 500-step cap, reported alongside a weighted-checkpoint partial score so progress short of completion is visible. Because grading inspects the end state, a plausible-looking trajectory that does not finish the job scores zero.
History and current status
XLANG Lab at the University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and elsewhere, released it in 2026 after OSWorld 1.0 was effectively solved. The board top is 20.6% full completion at a 54.8% partial score for Claude Opus 4.8 at maximum thinking as of June 2026. Anthropic reported Claude Opus 5 as ahead of every other model per unit cost at its July 2026 launch, without a board-verified figure.
What the score does not tell you
The absolute numbers are so low that ranking between models is fragile, and the partial-credit score and the completion score can tell different stories. Self-hosted websites and fixed input files make runs reproducible but also make the environment narrower than a real desktop. Cost is the other missing axis: a skilled human needs a median of roughly 1.6 hours per task, and a strong agent burns around 318 tool calls doing it worse.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports OSWorld 2.0, and how to read it
The XLANG board is the authoritative source, and it is academically run. Labs quote computer-use results at launch, sometimes as cost-efficiency claims rather than completion rates. Prefer the board figure, and note whether a claim refers to full completion or the weighted partial score, because they differ by more than thirty points.
Benchmarks to read alongside this one
OSWorld
Whether a multimodal agent can operate a real computer (desktop apps, file I/O, multi-app workflows) to complete open-ended tasks in a live virtual machine.
AndroidWorld
Whether an agent can operate a real Android phone to finish tasks across everyday apps.
Windows Agent Arena
Whether a multimodal agent can operate a full Windows desktop across the applications people actually use at work.
OSWorld 2.0: frequently asked questions
- What is OSWorld 2.0?
- OSWorld 2.0 is a 2026 benchmark of 108 long-horizon computer-use tasks across seven professional domains, run on a real desktop with 31 self-hosted websites and authentic input files. Grading is binary completion at a 500-step cap plus a weighted-checkpoint partial score.
- Why did OSWorld 2.0 replace OSWorld?
- Because agents reached about 83.5% on the original, so it no longer separated systems. OSWorld 2.0 uses far longer multi-application tasks, and scores dropped to roughly 20% full completion, which restored a great deal of headroom.
- What is the best OSWorld 2.0 score?
- About 20.6% full completion with a 54.8% partial score, recorded for Claude Opus 4.8 at maximum thinking as of June 2026. A skilled human takes a median of roughly 1.6 hours per task, which is the comparison that puts the number in context.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.