Tool use and MCP benchmark

MCP-Bench

Connects to 28 live MCP servers spanning about 250 tools. Part of the score uses an LLM judge (o4-mini).

Accenture

At a glance

What the agent must do
The agent must plan and run multi step, cross tool workflows on live MCP servers from fuzzy instructions that do not name the tools.
Size
104 tasks (56 single server, 30 two server, 18 three server)
Scoring
Overall score averaging schema understanding, LLM judged task completion, tool usage and planning
How to run it
Conda Python 3.10 env, install the bundled MCP servers with install.sh, then python run_benchmark.py --models <model>

Official sources

Other tool use and mcp benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.