Web browsing benchmark

WebArena

Widely used web agent benchmark; the repo's last news is from late 2024 and points to newer sibling projects such as TheAgentCompany and WebArena-Infinity. The environment must be reset to its initial state after a full run.

Carnegie Mellon University

At a glance

What the agent must do
The agent must complete long, realistic tasks on self hosted websites such as a shop, a forum, GitLab and a content manager.
Size
812 tasks
Scoring
Task success rate by programmatic functional correctness checks
How to run it
Self hosted websites via Docker or a prebuilt Amazon Machine Image; maintainers recommend the AgentLab framework for experiments

Official sources

Other web browsing benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.