Terminal benchmark

Terminal-Bench

Now a continuously versioned benchmark: 2.0 (Nov 2025), 2.1, 3.0 (74 tasks at launch, Jul 2026) and current 4.0 (Aug 2026), which recalibrated resources, fixed 19 tasks and removed 8. Results across major versions are not comparable, and the team published leaderboard integrity rules against cheating and reward hacking.

Stanford University and Laude Institute (Harbor team)

At a glance

What the agent must do
The agent must complete hard, realistic tasks inside a containerized terminal, from coding and system administration to science and ML work.
Size
Not stated in one number by the maintainers
Scoring
Resolution rate (% of tasks whose verifier passes), with 95% confidence intervals
How to run it
Harbor CLI: uv tool install 'harbor[modal]', then harbor run -d terminal-bench/terminal-bench@4.0.0 (some tasks need GPU sandboxes such as Modal)

Official sources

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.