Computer use benchmark

OSWorld 2.0

Tasks span 7 professional domains and 31 self hosted websites, with a median human time of about 1.6 hours and separate safety reports. Verified leaderboard entries are run by the maintainers on request.

XLANG Lab, University of Hong Kong and collaborators

At a glance

What the agent must do
The agent must carry out long, realistic computer workflows across apps and self hosted websites, often taking a human more than an hour.
Size
108 tasks
Scoring
Binary task completion within 500 steps, plus a partial score from checkpoints
How to run it
uv sync in the OSWorld-V2 repo; run on Docker (Linux with KVM) or AWS images

Official sources

Other computer use benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.