Tool use and MCP benchmark
Toolathlon (Tool Decathlon)
Toolathlon-Verified (June 2026) is the current release, with prompts, ground truths and evaluators reviewed. Environments start from realistic states such as real Canvas courses and spreadsheets.
HKUST NLP
At a glance
- What the agent must do
- The agent must finish multi step workflows across many real apps, such as email, calendars, Kubernetes and BigQuery, mostly through MCP servers.
- Size
- 108 tasks across 32 apps and 604 tools
- Scoring
- Pass@1 success rate from dedicated evaluation scripts (also Pass@3 and Pass^3)
- How to run it
- Public evaluation service with no local setup, or self hosted with uv and a Docker or Podman image
Official sources
Other tool use and mcp benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.