Sierra τ²-bench: a tool-using support agent plus a simulated user, including telecom where the user can call tools too.
unassessed
| Category | agentic |
|---|---|
| Subcategory | dual-control tool-agent-user customer-service simulations |
| Page status | active |
| Metric | pass^k (domain-specific success, then reliability across trials) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 375 |
| Dataset licence | MIT |
| Publisher | Sierra |
τ²-bench keeps the τ-bench loop — a policy-following agent talks to an LLM-simulated customer and calls domain APIs — and adds dual control. In telecom, the user also has tools that change a shared environment (a mocked phone), so the agent must instruct as well as act. Retail and airline remain constraint-satisfaction domains with a passive user. inspect_evals also ships banking_knowledge, a later knowledge-base domain from the same repo. English text and structured tools. This is not the original [tau_bench](tau_bench.md) retail/airline-only release.
Multi-turn Tool-Agent-User dialogue. Each trial needs an agent model and a user-simulator model. Official tau2 run uses four trials per task at temperature 0. inspect_evals tasks: tau2_airline, tau2_retail, tau2_telecom, tau2_banking. Banking retrieval_config is grep (default), full_kb, or no_knowledge; upstream leaderboard alltools is not ported.
No model card in ModelSpec reports this benchmark yet.