Tool use and MCP benchmark
MCP-Universe
Some tasks use dynamic evaluators that fetch live ground truth for time sensitive answers, so they need third party API keys such as Google Maps, GitHub and Notion.
Salesforce AI Research
At a glance
- What the agent must do
- The agent must solve hard tasks by calling real world MCP servers in areas like maps, GitHub, finance, 3D design, browser automation and web search.
- Size
- 231 tasks across 6 domains and 11 MCP servers
- Scoring
- Success rate from execution based evaluators (format, static and dynamic real time checks)
- How to run it
- Python 3.10+ and Docker; pip install requirements, add LLM and service API keys to .env, run per domain benchmark scripts
Official sources
Other tool use and mcp benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.