OSWorld 2.0
Also known as OSWorld-V2, OSWorld2.0
OSWorld 2.0 tests whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state. It exists because agents had reached 83.5% on the original OSWorld, and it shows how much of that progress was benchmark-specific: the best full-completion score here is about 20%.
| What it measures | Whether a computer-use agent can finish long-horizon professional work on a real desktop, coordinating several applications and leaving the machine in the correct final state. |
|---|---|
| Built by | XLANG Lab, University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and others, 2026 |
| Format | 108 long-horizon tasks across 7 professional domains and 21 sub-categories, using 31 self-hosted websites alongside desktop applications and authentic input files |
| Scoring metric | Binary completion at a 500-step cap, reported with a weighted-checkpoint partial score |
| Status | Active |
| Representative top score | 20.6% · Claude Opus 4.8 (max thinking) · read 2026-06 |
| Official leaderboard | osworld-v2.xlang.ai |
How OSWorld 2.0 works
The benchmark contains 108 long-horizon tasks across seven professional domains and 21 subcategories, using 31 self-hosted websites alongside real desktop applications and authentic input files. Grading is binary completion at a 500-step cap, reported alongside a weighted-checkpoint partial score so progress short of completion is visible. Because grading inspects the end state, a plausible-looking trajectory that does not finish the job scores zero.
History and current status
XLANG Lab at the University of Hong Kong, with collaborators at Columbia, UCSB, UCSD, Mila, Ohio State and elsewhere, released it in 2026 after OSWorld 1.0 was effectively solved. The board top is 20.6% full completion at a 54.8% partial score for Claude Opus 4.8 at maximum thinking as of June 2026. Anthropic reported Claude Opus 5 as ahead of every other model per unit cost at its July 2026 launch, without a board-verified figure.
What the score does not tell you
The absolute numbers are so low that ranking between models is fragile, and the partial-credit score and the completion score can tell different stories. Self-hosted websites and fixed input files make runs reproducible but also make the environment narrower than a real desktop. Cost is the other missing axis: a skilled human needs a median of roughly 1.6 hours per task, and a strong agent burns around 318 tool calls doing it worse.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Scored on those four axes, OSWorld 2.0 carries a concern score of 1 out of 8 and ranks 4 of 15, which puts it in the group worth quoting as it stands. See the reasoning behind that rating and how it compares with the rest of the field.
What OSWorld 2.0 scores actually mean
OSWorld 2.0 reports two numbers and both are needed. The leader reads 20.6% full completion and 54.8% on the weighted-checkpoint partial score. The distance between them is the story: agents routinely get most of the way through a long computer-use task and fail to finish it, and only the completion figure reflects work a person would not have to redo. Scale the numbers against the human baseline of roughly 1.6 hours median per task and about 318 tool calls for a strong agent, and a 20% completion rate on 108 long-horizon tasks describes a capability that is real but nowhere near unattended operation. The set exists because agents had reached 83.5% on OSWorld 1.0, so the low scores are deliberate headroom.
Who reports OSWorld 2.0, and how to read it
The XLANG board is the authoritative source, and it is academically run. Labs quote computer-use results at launch, sometimes as cost-efficiency claims rather than completion rates. Prefer the board figure, and note whether a claim refers to full completion or the weighted partial score, because they differ by more than thirty points.
When to weight OSWorld 2.0 in a model choice
Weight OSWorld 2.0 when the question is whether an agent can drive a desktop or a browser through a long professional task, and treat the full-completion column as the only one that maps to delivered work. It is the right benchmark for computer-use claims and the wrong one for code generation or reasoning. Because scores depend on the scaffolding and the step cap, confirm both before comparing two systems. Vendor claims of leading performance per unit cost that are not present on the official board should be labeled as unverified, since the board is the only place these runs are reproducible.
Benchmarks to read alongside this one
OSWorld
Whether a multimodal agent can operate a real computer (desktop apps, file I/O, multi-app workflows) to complete open-ended tasks in a live virtual machine.
AndroidWorld
Whether an agent can operate a real Android phone to finish tasks across everyday apps.
Windows Agent Arena
Whether a multimodal agent can operate a full Windows desktop across the applications people actually use at work.
OSWorld 2.0: frequently asked questions
- What is OSWorld 2.0?
- OSWorld 2.0 is a 2026 benchmark of 108 long-horizon computer-use tasks across seven professional domains, run on a real desktop with 31 self-hosted websites and authentic input files. Grading is binary completion at a 500-step cap plus a weighted-checkpoint partial score.
- Why did OSWorld 2.0 replace OSWorld?
- Because agents reached about 83.5% on the original, so it no longer separated systems. OSWorld 2.0 uses far longer multi-application tasks, and scores dropped to roughly 20% full completion, which restored a great deal of headroom.
- What is the best OSWorld 2.0 score?
- About 20.6% full completion with a 54.8% partial score, recorded for Claude Opus 4.8 at maximum thinking as of June 2026. A skilled human takes a median of roughly 1.6 hours per task, which is the comparison that puts the number in context.
Sources
- XLANG Lab et al. (2026). OSWorld 2.0 benchmark paper. arXiv preprint 2606.29537. arxiv.org/abs/2606.29537
- Xie, Zhang, Zhou, Zhou et al. (2024). OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv preprint 2404.07972, the predecessor set. arxiv.org/abs/2404.07972
- XLANG Lab. Official OSWorld 2.0 leaderboard. osworld-v2.xlang.ai
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →