τ-bench

Scores whether a tool-using agent can hold a policy-following conversation with a simulated customer and leave the backend in the right state.

Also known as: tau-bench, tau bench

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorytool-use customer-service agents
Page statussuperseded
Metricpass^k
Directionhigher_is_better
Unit%
Dataset size165
Dataset licenceMIT
PublisherSierra

What it measures

τ-bench asks a language agent to act as a customer-service representative in a retail or an airline domain: it chats with a user that is itself simulated by an LLM, calls domain-specific API tools to look up and change records, and must follow a written policy document. The user pursues its own goal and only reveals what it would naturally reveal, so the agent has to ask questions rather than execute a fixed call sequence.

Task format

multi-turn text dialogue interleaved with structured tool calls, against a simulated user model

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GLM 5.1Z.ai (Zhipu AI)70.62026-04
Claude Opus 4Anthropic68.22026-04
Claude Opus 4.6Anthropic68.22026-04
o3OpenAI65.52026-04
o3-deep-researchOpenAI65.52026-04
o3-miniOpenAI65.52026-04
o3-proOpenAI65.52026-04
Claude Opus 4.1Anthropic64.82026-04
Claude Opus 4.1 (latest)Anthropic64.82026-04
Claude Sonnet 4Anthropic61.52026-04
Claude Sonnet 4.5Anthropic61.52026-04
Claude Sonnet 4.5 (latest)Anthropic61.52026-04
GPT-5OpenAI58.52026-04
GPT-5 Chat (latest)OpenAI58.52026-04
GPT-5 MiniOpenAI58.52026-04
GPT-5 NanoOpenAI58.52026-04
GPT-5 ProOpenAI58.52026-04
GPT-5-CodexOpenAI58.52026-04
GPT-5.1OpenAI58.52026-04
GPT-5.1 ChatOpenAI58.52026-04
GPT-5.1 CodexOpenAI58.52026-04
GPT-5.1 Codex MaxOpenAI58.52026-04
GPT-5.1 Codex miniOpenAI58.52026-04
GPT-5.2OpenAI58.52026-04
GPT-5.2 ChatOpenAI58.52026-04
GPT-5.2 CodexOpenAI58.52026-04
GPT-5.2 ProOpenAI58.52026-04
GPT-5.3 Chat (latest)OpenAI58.52026-04
GPT-5.3 CodexOpenAI58.52026-04
GPT-5.3 Codex SparkOpenAI58.52026-04
GPT-5.4OpenAI58.52026-04
GPT-5.4 miniOpenAI58.52026-04
GPT-5.4 nanoOpenAI58.52026-04
GPT-5.4 ProOpenAI58.52026-04
Gemini 2.5 ProGoogle DeepMind58.22026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team55.82026-04
GPT-4.1OpenAI52.82026-04
Qwen 3 235B InstructCerebras50.22026-04
Qwen3 235B-A22BAlibaba / Qwen Team50.22026-04
DeepSeek R1DeepSeek48.82026-04
DeepSeek R1 0528DeepSeek48.82026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA48.82026-04
DeepSeek ReasonerDeepSeek48.82026-04
GPT-4oOpenAI42.52026-04
GPT-4o (2024-05-13)OpenAI42.52026-04
GPT-4o (2024-08-06)OpenAI42.52026-04
GPT-4o (2024-11-20)OpenAI42.52026-04
GPT-4o miniOpenAI42.52026-04
DeepSeek ChatDeepSeek41.22026-04
DeepSeek V3DeepSeek41.22026-04
DeepSeek V3 0324DeepSeek41.22026-04
DeepSeek V3.1DeepSeek41.22026-04
DeepSeek V3.2DeepSeek41.22026-04
DeepSeek V3.2 ExpDeepSeek41.22026-04

Data

This page as JSON · Edit on GitHub