General and professional work benchmark
METR Time Horizons
Built from RE-Bench, HCAST and shorter novel tasks; current version is Time Horizon 1.1. METR states measurements above 16 hours are unreliable with the current suite and that tasks are cleaner than real work.
METR
At a glance
- What the agent must do
- Agents attempt software, ML and cybersecurity tasks of known human duration, and the result is the task length at which they succeed half the time.
- Size
- Not stated in one number by the maintainers
- Scoring
- 50% (and 80%) task completion time horizon, in human expert minutes or hours
- How to run it
- See the official site
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.