Tool use and MCP benchmark

Berkeley Function Calling Leaderboard (BFCL)

Now at V4, which added agentic web search and memory categories. The paper appeared at ICML 2025 rather than on arXiv.

UC Berkeley (Gorilla LLM team)

At a glance

What the agent must do
The model must call functions and tools correctly, including multi turn use, memory, web search and avoiding hallucinated calls.
Size
Several thousand test entries; V4 includes 800 multi turn and 665 agentic entries
Scoring
Overall accuracy as a weighted average of category scores, checked by AST or state transition matching
How to run it
pip install bfcl-eval (package in berkeley-function-call-leaderboard)

Official sources

Other tool use and mcp benchmarks

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.