Coding benchmark
SWE-Lancer
The original SWELancer-Benchmark repo was archived in July 2025 and merged into openai/preparedness, now openai/frontier-evals. Only the Diamond split is public.
OpenAI
At a glance
- What the agent must do
- The agent must complete real paid Upwork software tasks, either implementing fixes and features or choosing the best proposal as a manager.
- Size
- Over 1,400 tasks worth $1 million in payouts; public subset is SWE-Lancer Diamond
- Scoring
- Tasks passed and dollars earned; coding tasks graded by end to end tests, manager tasks against the original manager's choice
- How to run it
- Unified Docker image; run from the SWE-Lancer directory of openai/frontier-evals
Official sources
Other coding benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.