About Tool-Agent-User Benchmark
Definition and scoring
- Organisation
- Shunyu Yao, Noah Shinn, Pedram Razavi, Karthik Narasimhan
- Category
- Agents
- Version
- 2024
- Direction
- higher is better
- Ranking use
- Reference
- Contamination risk
- Unknown
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints. Every genuine source row stays tied to its exact model label, configuration, benchmark version and evaluation system. The summary chart shows one best compatible score per canonical product; Score 2.0 admits only explicitly mapped, frozen protocols and keeps incompatible configurations separate.
Open benchmark source ↗