τ²-bench

Sierra τ²-bench: a tool-using support agent plus a simulated user, including telecom where the user can call tools too.

Also known as: tau2-bench, tau^2-bench, tau 2, inspect_evals/tau2

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorydual-control tool-agent-user customer-service simulations
Page statusactive
Metricpass^k (domain-specific success, then reliability across trials)
Directionhigher_is_better
Unit%
Dataset size375
Dataset licenceMIT
PublisherSierra

What it measures

τ²-bench keeps the τ-bench loop — a policy-following agent talks to an LLM-simulated customer and calls domain APIs — and adds dual control. In telecom, the user also has tools that change a shared environment (a mocked phone), so the agent must instruct as well as act. Retail and airline remain constraint-satisfaction domains with a passive user. inspect_evals also ships banking_knowledge, a later knowledge-base domain from the same repo. English text and structured tools. This is not the original [tau_bench](tau_bench.md) retail/airline-only release.

Task format

Multi-turn Tool-Agent-User dialogue. Each trial needs an agent model and a user-simulator model. Official tau2 run uses four trials per task at temperature 0. inspect_evals tasks: tau2_airline, tau2_retail, tau2_telecom, tau2_banking. Banking retrieval_config is grep (default), full_kb, or no_knowledge; upstream leaderboard alltools is not ported.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub