Toolathlon-Verified leaderboard
3 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 5 | #1 | Anthropic | 80.6% |
| Kimi K3 | #2 | Moonshot AI | 73.2% |
| Laguna S 2.1 | #3 | Poolside | 49.7% |
Agents · Benchmark profile
A verified tool-use benchmark variant for completing multi-step workflows with external tools.
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on Toolathlon-Verified
Claude Opus 5 leads at 80.6%, followed by Kimi K3 (73.2%) and Laguna S 2.1 (49.7%).
3 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| Claude Opus 5 | #1 | Anthropic | 80.6% |
| Kimi K3 | #2 | Moonshot AI | 73.2% |
| Laguna S 2.1 | #3 | Poolside | 49.7% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Claude Opus 5 claude-opus-5 | Anthropic | closed | Estimated reference | 80.6% |
| #2 | Kimi K3 kimi-k3 | Moonshot AI | closed | Reference only | 73.2% |
| #3 | Laguna S 2.1 laguna-s-2-1 | Poolside | open | Estimated reference | 49.7% |
The top of this snapshot is led by Claude Opus 5 at 80.6%; third place is 30.9 points behind. The top-3 spread is 30.9 points.
About Toolathlon-Verified
A verified tool-use benchmark variant for completing multi-step workflows with external tools. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A verified tool-use benchmark variant for completing multi-step workflows with external tools.
Claude Opus 5 by Anthropic currently leads with 80.6%.
3 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related