Tool use and MCP benchmark

MCP-Universe

Some tasks use dynamic evaluators that fetch live ground truth for time sensitive answers, so they need third party API keys such as Google Maps, GitHub and Notion.

Salesforce AI Research

At a glance

What the agent must do
The agent must solve hard tasks by calling real world MCP servers in areas like maps, GitHub, finance, 3D design, browser automation and web search.
Size
231 tasks across 6 domains and 11 MCP servers
Scoring
Success rate from execution based evaluators (format, static and dynamic real time checks)
How to run it
Python 3.10+ and Docker; pip install requirements, add LLM and service API keys to .env, run per domain benchmark scripts

Official sources

Other tool use and mcp benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.