WebArena leaderboard
1 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| Muse Spark 1.1 | #1 | Meta | 69% |
Agents · Benchmark profile
A benchmark of autonomous agents completing realistic tasks on self-hosted web applications.
Data verified 14 Jul 2026 · Methodology 1.8.0
1 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| Muse Spark 1.1 | #1 | Meta | 69% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | Muse Spark 1.1 muse-spark-1-1 | Meta | closed | Reference only | 69% |
About WebArena
A benchmark of autonomous agents completing realistic tasks on self-hosted web applications. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A benchmark of autonomous agents completing realistic tasks on self-hosted web applications.
Muse Spark 1.1 by Meta currently leads with 69%.
1 model in the LuminaBench cohort have a qualifying score on this benchmark.
No. This benchmark is display-only and does not enter the overall Lumina composite.
Related