Coding benchmark
SWE-bench
The original benchmark built from real GitHub issues and pull requests. The site also lists a 300 task Lite subset for cheaper runs and a 300 task Multilingual set covering 9 languages.
Princeton University (SWE-bench team)
At a glance
- What the agent must do
- Given a real Python repository and a GitHub issue, the agent must edit the code so the issue is fixed and the hidden tests pass.
- Size
- 2,294 tasks from 12 Python repositories
- Scoring
- % of issues resolved (tests pass)
- How to run it
- Docker harness: clone the repo, pip install -e ., then run swebench eval with a predictions file
Official sources
Other coding benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.