Coding benchmark

SWE-Lancer

The original SWELancer-Benchmark repo was archived in July 2025 and merged into openai/preparedness, now openai/frontier-evals. Only the Diamond split is public.

OpenAI

At a glance

What the agent must do
The agent must complete real paid Upwork software tasks, either implementing fixes and features or choosing the best proposal as a manager.
Size
Over 1,400 tasks worth $1 million in payouts; public subset is SWE-Lancer Diamond
Scoring
Tasks passed and dollars earned; coding tasks graded by end to end tests, manager tasks against the original manager's choice
How to run it
Unified Docker image; run from the SWE-Lancer directory of openai/frontier-evals

Official sources

Other coding benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.