Customer service benchmark
tau-bench (tau2-bench, tau3)
Domains are retail and airline (original), telecom with dual control where the user also acts (tau2, arXiv 2506.07982), and banking knowledge retrieval and voice (tau3 updates, 2026). The repo warns results before version 1.0.1 are not comparable with later results.
Sierra
At a glance
- What the agent must do
- The agent must serve a simulated customer over a conversation, using domain tools and following a written company policy.
- Size
- Not stated in one number by the maintainers
- Scoring
- pass^k reliability across k trials, judged on the final database state; leaderboard reports Pass^1
- How to run it
- Clone tau2-bench, uv sync, then tau2 run --domain <airline|retail|telecom|banking_knowledge> --agent-llm <model> --user-llm <model>
Official sources
Run evals on your own agent
Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.
Get the weekly agent stack update
New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.
Reviewed Oct 9, 2026.