Computer use benchmark
OSWorld-Verified
OSWorld-Verified (July 2025) fixed 300+ community reported issues such as changed websites, ambiguous instructions and fragile checkers, so older OSWorld scores should not be compared directly. OSWorld 2.0 was released in June 2026 as the newer version.
XLANG Lab, University of Hong Kong (with Salesforce Research, CMU, University of Waterloo)
At a glance
- What the agent must do
- The agent must complete open ended tasks on a real desktop computer across web and desktop apps, operating the screen with mouse and keyboard.
- Size
- 369 tasks (361 if the 8 Google Drive tasks are excluded)
- Scoring
- Task success rate via execution based evaluation scripts
- How to run it
- VM based environment; VMware, Docker or AWS, with AWS parallelization bringing a full run to about an hour
Official sources
Other computer use benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.