Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.
Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.
Which model scores highest on Vibe Code Bench v1.1?
Claude Opus 4.7 by Anthropic currently leads with 71.0%.
How many models are evaluated on Vibe Code Bench v1.1?
41 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.