Berkeley Function-Calling Leaderboard leaderboard
2 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Qwen3.7-Max | #1 | Alibaba Cloud | 75% |
| Qwen3.7-Plus | #2 | Alibaba Cloud | 72.9% |
Agents · Benchmark profile
An evaluation suite for function selection, argument construction and tool-use behaviour.
Data verified 14 Jul 2026 · Methodology 1.8.0
Benchmark score on Berkeley Function-Calling Leaderboard
Qwen3.7-Max leads at 75%, followed by Qwen3.7-Plus (72.9%).
2 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| Qwen3.7-Max | #1 | Alibaba Cloud | 75% |
| Qwen3.7-Plus | #2 | Alibaba Cloud | 72.9% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Qwen3.7-Max qwen3.7-max | Alibaba Cloud | closed | Reference only | 75% |
| #2 | Qwen3.7-Plus qwen3.7-plus | Alibaba Cloud | closed | Reference only | 72.9% |
About Berkeley Function-Calling Leaderboard
An evaluation suite for function selection, argument construction and tool-use behaviour. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
An evaluation suite for function selection, argument construction and tool-use behaviour.
Qwen3.7-Max by Alibaba Cloud currently leads with 75%.
2 models in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related