Computer use benchmark

OSWorld-Verified

OSWorld-Verified (July 2025) fixed 300+ community reported issues such as changed websites, ambiguous instructions and fragile checkers, so older OSWorld scores should not be compared directly. OSWorld 2.0 was released in June 2026 as the newer version.

XLANG Lab, University of Hong Kong (with Salesforce Research, CMU, University of Waterloo)

At a glance

What the agent must do
The agent must complete open ended tasks on a real desktop computer across web and desktop apps, operating the screen with mouse and keyboard.
Size
369 tasks (361 if the 8 Google Drive tasks are excluded)
Scoring
Task success rate via execution based evaluation scripts
How to run it
VM based environment; VMware, Docker or AWS, with AWS parallelization bringing a full run to about an hour

Official sources

Other computer use benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.