Web browsing benchmark
WebArena
Widely used web agent benchmark; the repo's last news is from late 2024 and points to newer sibling projects such as TheAgentCompany and WebArena-Infinity. The environment must be reset to its initial state after a full run.
Carnegie Mellon University
At a glance
- What the agent must do
- The agent must complete long, realistic tasks on self hosted websites such as a shop, a forum, GitLab and a content manager.
- Size
- 812 tasks
- Scoring
- Task success rate by programmatic functional correctness checks
- How to run it
- Self hosted websites via Docker or a prebuilt Amazon Machine Image; maintainers recommend the AgentLab framework for experiments
Official sources
Other web browsing benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.