τ-bench evaluates whether tool-using agents complete user goals while following domain policies across stateful conversations.
unassessed
| Category | agentic |
|---|---|
| Metric | pass^k |
| Direction | higher_is_better |
| Unit | fraction |
| Publisher | Sierra |
τ-bench evaluates an agent that talks with a simulated user and calls domain-specific APIs. The benchmark checks both task completion and adherence to written policies in dynamic retail and airline domains.
Multi-turn text conversations with tool calls and a final database state.
No model card in ModelSpec reports this benchmark yet.