Coding benchmark

SWE-bench

The original benchmark built from real GitHub issues and pull requests. The site also lists a 300 task Lite subset for cheaper runs and a 300 task Multilingual set covering 9 languages.

Princeton University (SWE-bench team)

At a glance

What the agent must do
Given a real Python repository and a GitHub issue, the agent must edit the code so the issue is fixed and the hidden tests pass.
Size
2,294 tasks from 12 Python repositories
Scoring
% of issues resolved (tests pass)
How to run it
Docker harness: clone the repo, pip install -e ., then run swebench eval with a predictions file

Official sources

Other coding benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.