Coding benchmark

SWE-bench Verified

Created in 2024 because some original SWE-bench tasks had overly specific tests, underspecified issues or unreliable environments. In February 2026 OpenAI stopped reporting it, citing flawed tests in an audit and evidence that frontier models had memorized the tasks, and recommended SWE-bench Pro instead.

OpenAI with the SWE-bench authors (Princeton)

At a glance

What the agent must do
The agent must resolve real GitHub issues in Python repositories, using a subset of SWE-bench that human engineers screened for fair tests and clear issue text.
Size
500 tasks
Scoring
% of issues resolved (tests pass)
How to run it
Docker harness via the SWE-bench repo: swebench eval verified -p <predictions> --run-id <id>

Official sources

Other coding benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.