Coding benchmark

SWE-Bench Pro

Designed to reduce contamination by using copyleft licensed public repos and private startup codebases; held out and commercial problems are not released. OpenAI recommends its public split in place of SWE-bench Verified, while calling it imperfect.

Scale AI

At a glance

What the agent must do
The agent must complete long horizon bug fixes and feature work in real business, B2B and developer tool codebases.
Size
1,865 tasks (731 public, 858 held out, 276 commercial)
Scoring
% of tasks resolved (fail to pass and pass to pass tests)
How to run it
Docker based; Modal recommended, local Docker in beta, via swe_bench_pro_eval.py

Official sources

Other coding benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.