Terminal benchmark
Terminal-Bench
Now a continuously versioned benchmark: 2.0 (Nov 2025), 2.1, 3.0 (74 tasks at launch, Jul 2026) and current 4.0 (Aug 2026), which recalibrated resources, fixed 19 tasks and removed 8. Results across major versions are not comparable, and the team published leaderboard integrity rules against cheating and reward hacking.
Stanford University and Laude Institute (Harbor team)
At a glance
- What the agent must do
- The agent must complete hard, realistic tasks inside a containerized terminal, from coding and system administration to science and ML work.
- Size
- Not stated in one number by the maintainers
- Scoring
- Resolution rate (% of tasks whose verifier passes), with 95% confidence intervals
- How to run it
- Harbor CLI: uv tool install 'harbor[modal]', then harbor run -d terminal-bench/terminal-bench@4.0.0 (some tasks need GPU sandboxes such as Modal)
Official sources
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.