Score is LuminaBench's central estimate of a model's broad capability across agents, coding, reasoning, research, multimodal understanding and mathematics. A higher Score means stronger performance across the complete weighted mix—not simply one impressive benchmark result. Overall and category positions use this central Score. The P10–P90 range remains visible as uncertainty, but its optimistic upper bound never determines order.
Each direct result is converted to a percentile inside that benchmark's own source snapshot, after applying its higher-is-better or lower-is-better direction. This makes an accuracy percentage, Elo score and task count comparable without averaging their raw scales.
Formula
pᵦ(x) = (count below x + ½(count tied with x − 1)) ÷ (nᵦ − 1)
Dbenchmark = 15 + 70pᵦRelated benchmark snapshots are combined inside one reviewed family, then the available families use their fixed published weights. Coding gives SWE-Bench Pro and LiveCodeBench equal direct-family weight; display-only specialist tests contribute nothing.
Formula
Dfamily = mean(Dbenchmark within family)
Dcategory = Σ(wf × Dfamily) ÷ Σ(available wf)BenchLM, Artificial Analysis and raw benchmark families use different scales. Each source is therefore converted to a percentile inside its own dated snapshot before sources are combined. A source's scale cannot dominate merely because its numbers are larger.
Formula
pₛ(x) = (count below x + ½(count tied with x − 1)) ÷ (nₛ − 1)
Eₛ = 15 + 70pₛ
E = mean(Eₛ across available independent sources)The versioned hierarchy record names reviewed capability classes and predecessor links. It is a transparent prior for sparse evidence, not a benchmark result. Specialist models receive no general hierarchy advantage outside the category they demonstrate.
Formula
H ∈ {82, 77, 72, 65, 58, 50}
Frontier flagship 82 · Frontier 77 · Current efficient 72
Established 65 · Legacy 58 · Specialist 50D is direct benchmark evidence, E is the current source-normalised external signal and H is the reviewed hierarchy signal. Unavailable terms are omitted and the remaining weights are re-normalised. A category with no evidence is clearly Estimated from H with a wide interval; it never inherits Overall Score.
Formula
Supported = 0.55D + 0.35E + 0.10H
Provisional = 0.35D + 0.45E + 0.20H
Estimated = 0.15D + 0.55E + 0.30H
No direct evidence = 0.65E + 0.35H
No category evidence = 1.00HSuccessors normally remain above predecessors. An inversion is allowed only when the published strong-evidence gate passes, or when a dated exception is recorded. Reviewed ordering anchors use the same minimum-correction pass. Raw Score, lineage adjustment and review adjustment remain separate in every ranking record.
Formula
δ* = arg min Σδm²
subject to Ssuccessor + δsuccessor ≥ Spredecessor + δpredecessor + 0.1
|δlineage + δreview| ≤ 6
unsatisfied constraint within cap → generation failsOverall and every category sort by the final central Ranking Score. The top 100 receive exact ordinal positions. Coding uses SWE-Bench Pro and LiveCodeBench as equal direct families; superseded specialist tests remain display-only. Research prioritises browsing, evidence gathering and synthesis, while AA-LCR belongs to long-context Reasoning.
Formula
Raw Score = wD·D + wE·E + wH·H
Ranking Score = Raw Score + δlineage + δreview
Rank = descending(Ranking Score)
80% interval = displayed uncertainty onlyThe final central capability estimate: Raw Score plus separately disclosed lineage and review corrections.
Supported, Provisional, Estimated or Insufficient, based on frozen evidence gates.
The central 80% Score and rank range from 1,000 deterministic bootstrap draws.
The weights are published before scores are generated and total 100%. Every family retains its weight even when a model has no direct result.
| Category | Weight | Benchmark families |
|---|---|---|
| Agents | 22% | Terminal-Bench · Professional work · Tool and computer use · Other agent tasks |
| Coding | 24% | SWE-Bench Pro · CursorBench (display only) · LiveCodeBench · Other coding (display only) |
| Reasoning | 16% | Humanity’s Last Exam · GPQA · ARC-AGI and general reasoning |
| Research | 14% | Web research and source finding · Professional research and synthesis · Factual research support |
| Multimodal | 12% | MMMU · CharXiv · Other multimodal |
| Mathematics | 12% | AIME and olympiad maths · FrontierMath · HMMT and other maths |
Methodology 1.8.0 uses the dated Lumina publisher ranking review record from 2026-07-28. It publishes 27 explicit model-class or predecessor entries, reviewed Overall and Coding anchors, the permitted exception record and its provenance. These judgements calibrate sparse evidence and never masquerade as direct benchmark results.
The combined lineage and review correction is bounded to ±6.0 Score points per model in Overall and each category. Generation stops rather than silently exceeding that cap.
Missing benchmarks stay marked Estimated. A category with external evidence uses E and H; one with no category evidence uses H alone and receives a wide interval. Neither route borrows Overall Score.
Speed, latency, context length, API price, task cost, composite confidence, advertising, sponsorship, News and commercial relationships remain separate from Score.
See Score, Coverage and Confidence together in the published ranking.
Open leaderboard →