Graphwalks BFS 0K-128K leaderboard
1 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| MAI-Thinking-1 | #1 | Microsoft | 90% |
Reasoning · Benchmark profile
Long-context graph traversal benchmark using breadth-first search tasks.
Data verified 27 Jul 2026 · Methodology 1.8.0
1 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| MAI-Thinking-1 | #1 | Microsoft | 90% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | MAI-Thinking-1 mai-thinking-1 | Microsoft | closed | Estimated reference | 90% |
About Graphwalks BFS 0K-128K
Long-context graph traversal benchmark using breadth-first search tasks. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
Long-context graph traversal benchmark using breadth-first search tasks.
MAI-Thinking-1 by Microsoft currently leads with 90%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
Related