Web browsing benchmark

BrowseComp

Includes a canary string and asks people not to post examples online to limit contamination. OpenAI notes short single answer questions may not reflect real user queries, and simple-evals is no longer updated though it keeps BrowseComp as a reference.

OpenAI

At a glance

What the agent must do
The agent must persistently browse the internet to find hard to locate, entangled facts with short, checkable answers.
Size
1,266 questions
Scoring
Accuracy against short reference answers
How to run it
Reference implementation browsecomp_eval.py in simple-evals

Official sources

Other web browsing benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.