General and professional work benchmark
Harbor-Index
Released June 2026, distilled from 6,627 candidate tasks across 54 benchmarks, keeping tasks that are hard but not broken. Meant as a cheap, high signal alternative to running every benchmark.
Harbor and Terminal-Bench team (Stanford University, Laude Institute)
At a glance
- What the agent must do
- The agent must solve a compact, hard mix of tasks drawn from many existing agent benchmarks, run through one standard harness.
- Size
- 82 tasks
- Scoring
- Pass rate, plotted against run cost
- How to run it
- Runs in the Harbor framework; tasks are on the Harbor Hub
Official sources
Other general and professional work benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.