Computer use benchmark
OSWorld 2.0
Tasks span 7 professional domains and 31 self hosted websites, with a median human time of about 1.6 hours and separate safety reports. Verified leaderboard entries are run by the maintainers on request.
XLANG Lab, University of Hong Kong and collaborators
At a glance
- What the agent must do
- The agent must carry out long, realistic computer workflows across apps and self hosted websites, often taking a human more than an hour.
- Size
- 108 tasks
- Scoring
- Binary task completion within 500 steps, plus a partial score from checkpoints
- How to run it
- uv sync in the OSWorld-V2 repo; run on Docker (Linux with KVM) or AWS images
Official sources
Other computer use benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.