Tool use and MCP benchmark
Berkeley Function Calling Leaderboard (BFCL)
Now at V4, which added agentic web search and memory categories. The paper appeared at ICML 2025 rather than on arXiv.
UC Berkeley (Gorilla LLM team)
At a glance
- What the agent must do
- The model must call functions and tools correctly, including multi turn use, memory, web search and avoiding hallucinated calls.
- Size
- Several thousand test entries; V4 includes 800 multi turn and 665 agentic entries
- Scoring
- Overall accuracy as a weighted average of category scores, checked by AST or state transition matching
- How to run it
- pip install bfcl-eval (package in berkeley-function-call-leaderboard)
Official sources
Other tool use and mcp benchmarks
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.