Harbor Framework (creators of Terminal-Bench)

Harbor

Framework for running and grading agents on benchmark or custom tasks in parallel sandboxes, from the Terminal-Bench team.

Apache-2.0Pythonv0.24.0

What it supports

  • Evaluates multi-step agents and tool calls: supportedEvaluates multi-step agents and tool callsDocs
  • LLM as a judge scoring: supportedLLM as a judge scoringDocs
  • Tracing: supportedTracingDocs
  • ?Runs in CI: not yet verifiedRuns in CI

Install

pip install harbor

Links

Other eval frameworks

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.