Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.
Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.
Which model scores highest on FrontierCode 1.1 Main?
Claude Fable 5 by Anthropic currently leads with 53.5%.
How many models are evaluated on FrontierCode 1.1 Main?
10 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.