τ-bench

τ-bench evaluates whether tool-using agents complete user goals while following domain policies across stateful conversations.

Also known as: tau-bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Metricpass^k
Directionhigher_is_better
Unitfraction
PublisherSierra

What it measures

τ-bench evaluates an agent that talks with a simulated user and calls domain-specific APIs. The benchmark checks both task completion and adherence to written policies in dynamic retail and airline domains.

Task format

Multi-turn text conversations with tool calls and a final database state.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub