An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
Which model scores highest on ResearchClawBench?
Claude Opus 4.8 by Anthropic currently leads with 21.1%.
How many models are evaluated on ResearchClawBench?
20 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.