Reasoning
Mathematical and scientific problem solving.
Model profile / Meta
Llama 4 Scout is a open weight model variant recorded in the BenchLM public dataset.
Four areas of evidence
Mathematical and scientific problem solving.
Software engineering and programming.
Factual reliability and instruction following.
Tool use, workflows and professional tasks.
Current index snapshot · 2026-09-13. Each category uses its own scale; scores across categories are not directly comparable. Evidence counts describe the retained Capabilities inputs; a family can contain several tasks or configurations. A point lead does not establish superiority.
Performance in context
Llama 4 Scout is highlighted wherever a compatible measurement is available.
Llama 4 Scout: First-party page did not expose a verifiable checkpoint-specific token price; no AA/provider-median substitute
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Lumina Capabilities Index
The current Capabilities cohort, with this model highlighted.
Chart loads as you explore
Explore the detail
Score uses the admitted direct native evidence under the selected method.
| Benchmark | Result | Published configuration | Source |
|---|---|---|---|
| AA Long Context Reasoningstandard | 30.333% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis AA-LCR evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| CritPtstandard | 0% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis CritPt evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GDPval-AA v2v2 | 0% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis GDPval-AA v2 | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| GPQA Diamonddiamond | 58.687% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis GPQA Diamond evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Humanity's Last Examtext-only current | 3.78% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis text-only HLE evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| SciCode2024 | 17.014% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis SciCode evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
| τ³-Banking3 | 3.299% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis tau3-Banking evaluation | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Terminal-Bench2.1 | 3.745% | Llama 4 Scout (Artificial Analysis completed independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
| Artificial Analysis Agentic Index2026 | 1.1 index | Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
| Artificial Analysis Agentic Index2026 | 1.1 index | Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
No specialist task observations retained for this model.