SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.
Reference2025Active19 models
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
About SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines
Definition and scoring
Organisation
Xiaoxuan Du, Yao Yao, Kexin Ma, Bowen Wang, Tianyu Zheng, Kaiyan Zhu, Yiming Zhang, Yutao Zhu, Jiawei Zhou, Jingren Zhou
Category
Knowledge
Version
2025
Direction
higher is better
Ranking use
Reference
Contamination risk
Unknown
An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines measure?
An expanded version of GPQA that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines, providing comprehensive coverage of academic domains.
Which model scores highest on SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines?
Claude Opus 4.6 by Anthropic currently leads with 95%.
How many models are evaluated on SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines?
19 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.