τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
Reference2026Active10 models
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on τ³-Bench Tool-Agent-User Evaluation
τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does τ³-Bench Tool-Agent-User Evaluation measure?
τ³-bench is the current evolution of Sierra's tool-agent-user framework, adding corrected task releases and newer knowledge and voice evaluation modes alongside airline, retail, and telecom.
Which model scores highest on τ³-Bench Tool-Agent-User Evaluation?
Mistral Medium 3.5 128B by Mistral AI currently leads with 91.4%.
How many models are evaluated on τ³-Bench Tool-Agent-User Evaluation?
10 models in the LuminaBench cohort have a qualifying score on this benchmark.