Web browsing benchmark
BrowseComp
Includes a canary string and asks people not to post examples online to limit contamination. OpenAI notes short single answer questions may not reflect real user queries, and simple-evals is no longer updated though it keeps BrowseComp as a reference.
OpenAI
At a glance
- What the agent must do
- The agent must persistently browse the internet to find hard to locate, entangled facts with short, checkable answers.
- Size
- 1,266 questions
- Scoring
- Accuracy against short reference answers
- How to run it
- Reference implementation browsecomp_eval.py in simple-evals
Official sources
Other web browsing benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.