A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.
Reference2026Active54 models
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on Gert Labs Composite Game Benchmark
A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does Gert Labs Composite Game Benchmark measure?
A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.
Which model scores highest on Gert Labs Composite Game Benchmark?
Claude Opus 4.8 by Anthropic currently leads with 73.0%.
How many models are evaluated on Gert Labs Composite Game Benchmark?
54 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.