A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
Reference2023Active70 models
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on Graduate-Level Google-Proof Q&A
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, Samuel R. Bowman
Category
Knowledge
Version
2023
Direction
higher is better
Ranking use
Reference
Contamination risk
Unknown
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does Graduate-Level Google-Proof Q&A measure?
A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Designed to be difficult even for skilled non-experts with access to Google.
Which model scores highest on Graduate-Level Google-Proof Q&A?
Sakana Fugu by Sakana AI currently leads with 95.5%.
How many models are evaluated on Graduate-Level Google-Proof Q&A?
70 models in the LuminaBench cohort have a qualifying score on this benchmark.