Capital & Compute
Benchmark· Agents, tool use & computer use· Checked 2026-06-29

tau-bench

Also known as τ-bench, tau2-bench, τ2-bench

tau-bench measures whether a tool-using agent completes customer-service tasks reliably, not just occasionally. Its distinguishing feature is the metric: pass^k, the probability that an agent succeeds in all k independent attempts. That reframes the question from "can it do this" to "can it do this every time," which is the question that actually decides whether an agent can be deployed.

Key facts about the tau-bench benchmark
What it measuresWhether a tool-using agent can reliably complete customer-service tasks over multi-turn conversations with a simulated user while obeying domain policies.
Built bySierra (Yao, Shinn, Narasimhan et al.), 2024
Format165 tasks in v1 (115 retail, 50 airline) as dynamic dialogues with a simulated user plus domain APIs; later versions add telecom and banking
Scoring metricpass^k: the probability an agent succeeds across all k independent trials (reliability, not just average success)
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard

How tau-bench works

Version 1 contains 165 tasks, 115 in a retail domain and 50 in an airline domain, later extended to telecom and banking. Each task is a dynamic multi-turn conversation with a simulated user, plus a set of domain APIs the agent must call, and a set of domain policies it must not violate. Success requires the correct final database state and policy compliance, and the headline reliability figure is pass^k rather than average success.

History and current status

Sierra introduced tau-bench in 2024, at a point when agent demos were impressive and agent products were not shipping. It became the standard citation for the reliability gap, and later versions broadened the domains. The tau2-bench line continues the design with dynamic environments derived from further real service domains.

What the score does not tell you

The simulated user is itself a language model, so part of what is being measured is the interaction between two models rather than an agent facing a person. Domain coverage is narrow, and policy compliance is defined by the benchmark authors rather than by any real operator. None of that undermines the central finding, which is that pass^1 and pass^8 differ dramatically for every model tested.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports tau-bench, and how to read it

Labs quote tau-bench selectively, and almost always as pass^1, which is the flattering number. The reliability story lives in the gap between pass^1 and pass^k. When a launch reports only single-attempt success on an agentic benchmark, that omission is the finding.

2024
First released
Sierra (Yao, Shinn, Narasimhan et al.)
Active
Status today
As of July 27, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

tau-bench: frequently asked questions

What is tau-bench?
tau-bench is a 2024 benchmark from Sierra that tests whether a tool-using agent can complete customer-service tasks over multi-turn conversations with a simulated user while obeying domain policies. Version 1 has 165 tasks across retail and airline domains.
What is pass^k?
pass^k is the probability that an agent succeeds on all k independent attempts at the same task. Unlike pass@k, which rewards succeeding at least once, pass^k measures consistency. It is the metric that exposes agents which work in a demo but not in production.
Which benchmarks measure AI agent reliability?
tau-bench is the primary one, because pass^k measures consistency directly. Terminal-Bench, OSWorld 2.0, AutomationBench, GAIA, AgentBench and WebArena all contribute, and Vending-Bench targets long-horizon coherence specifically.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory