Coding benchmark
SWE-bench Verified
Created in 2024 because some original SWE-bench tasks had overly specific tests, underspecified issues or unreliable environments. In February 2026 OpenAI stopped reporting it, citing flawed tests in an audit and evidence that frontier models had memorized the tasks, and recommended SWE-bench Pro instead.
OpenAI with the SWE-bench authors (Princeton)
At a glance
- What the agent must do
- The agent must resolve real GitHub issues in Python repositories, using a subset of SWE-bench that human engineers screened for fair tests and clear issue text.
- Size
- 500 tasks
- Scoring
- % of issues resolved (tests pass)
- How to run it
- Docker harness via the SWE-bench repo: swebench eval verified -p <predictions> --run-id <id>
Official sources
Other coding benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.