Tool use and MCP benchmark
MCP-Bench
Connects to 28 live MCP servers spanning about 250 tools. Part of the score uses an LLM judge (o4-mini).
Accenture
At a glance
- What the agent must do
- The agent must plan and run multi step, cross tool workflows on live MCP servers from fuzzy instructions that do not name the tools.
- Size
- 104 tasks (56 single server, 30 two server, 18 three server)
- Scoring
- Overall score averaging schema understanding, LLM judged task completion, tool usage and planning
- How to run it
- Conda Python 3.10 env, install the bundled MCP servers with install.sh, then python run_benchmark.py --models <model>
Official sources
Other tool use and mcp benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.