Tool use and MCP benchmark

AppWorld

Tests also check for collateral damage, meaning unexpected changes to app state. Test splits are test_normal and test_challenge.

Stony Brook University NLP

At a glance

What the agent must do
The agent must write interactive code against 457 APIs across 9 simulated everyday apps to complete tasks for simulated users.
Size
750 tasks
Scoring
Task Goal Completion (% of tasks passing all state based unit tests) and Scenario Goal Completion
How to run it
pip install appworld, then appworld install; APIs also available via an AppWorld MCP server

Official sources

Other tool use and mcp benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.