Customer service benchmark

tau-bench (tau2-bench, tau3)

Domains are retail and airline (original), telecom with dual control where the user also acts (tau2, arXiv 2506.07982), and banking knowledge retrieval and voice (tau3 updates, 2026). The repo warns results before version 1.0.1 are not comparable with later results.

Sierra

At a glance

What the agent must do
The agent must serve a simulated customer over a conversation, using domain tools and following a written company policy.
Size
Not stated in one number by the maintainers
Scoring
pass^k reliability across k trials, judged on the final database state; leaderboard reports Pass^1
How to run it
Clone tau2-bench, uv sync, then tau2 run --domain <airline|retail|telecom|banking_knowledge> --agent-llm <model> --user-llm <model>

Official sources

Run evals on your own agent

Benchmarks compare models and agents in general. To test your own agent, see the eval frameworks comparison.

Get the weekly agent stack update

New official MCP servers, spec changes and harness releases, checked against the source. One email a week, no fluff.

Reviewed Oct 9, 2026.