General and professional work benchmark

Harbor-Index

Released June 2026, distilled from 6,627 candidate tasks across 54 benchmarks, keeping tasks that are hard but not broken. Meant as a cheap, high signal alternative to running every benchmark.

Harbor and Terminal-Bench team (Stanford University, Laude Institute)

At a glance

What the agent must do
The agent must solve a compact, hard mix of tasks drawn from many existing agent benchmarks, run through one standard harness.
Size
82 tasks
Scoring
Pass rate, plotted against run cost
How to run it
Runs in the Harbor framework; tasks are on the Harbor Hub

Official sources

Other general and professional work benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.