Tool use and MCP benchmark
AppWorld
Tests also check for collateral damage, meaning unexpected changes to app state. Test splits are test_normal and test_challenge.
Stony Brook University NLP
At a glance
- What the agent must do
- The agent must write interactive code against 457 APIs across 9 simulated everyday apps to complete tasks for simulated users.
- Size
- 750 tasks
- Scoring
- Task Goal Completion (% of tasks passing all state based unit tests) and Scenario Goal Completion
- How to run it
- pip install appworld, then appworld install; APIs also available via an AppWorld MCP server
Official sources
Other tool use and mcp benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.