Coding benchmark
SWE-Bench Pro
Designed to reduce contamination by using copyleft licensed public repos and private startup codebases; held out and commercial problems are not released. OpenAI recommends its public split in place of SWE-bench Verified, while calling it imperfect.
Scale AI
At a glance
- What the agent must do
- The agent must complete long horizon bug fixes and feature work in real business, B2B and developer tool codebases.
- Size
- 1,865 tasks (731 public, 858 held out, 276 commercial)
- Scoring
- % of tasks resolved (fail to pass and pass to pass tests)
- How to run it
- Docker based; Modal recommended, local Docker in beta, via swe_bench_pro_eval.py
Official sources
Other coding benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.