SimpleQA leaderboard
2 ranked models · higher is better · labels show rank and score
View accessible chart data
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro | #1 | DeepSeek | 45% |
| DeepSeek V4 Flash | #2 | DeepSeek | 23.1% |
Research · Benchmark profile
A short-answer factuality benchmark designed to measure correctness on fact-seeking questions.
Data verified 15 Jul 2026 · Methodology 1.8.0
Benchmark score on SimpleQA
DeepSeek V4 Pro leads at 45%, followed by DeepSeek V4 Flash (23.1%).
2 ranked models · higher is better · labels show rank and score
| Model | Rank | Provider | Score |
|---|---|---|---|
| DeepSeek V4 Pro | #1 | DeepSeek | 45% |
| DeepSeek V4 Flash | #2 | DeepSeek | 23.1% |
One best score per model · higher is better
| Rank | Model | Provider | License | Evidence use | Score |
|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro deepseek-v4-pro | DeepSeek | open | Ranking eligible | 45% |
| #2 | DeepSeek V4 Flash deepseek-v4-flash | DeepSeek | open | Ranking eligible | 23.1% |
About SimpleQA
A short-answer factuality benchmark designed to measure correctness on fact-seeking questions. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Open benchmark source ↗FAQ
A short-answer factuality benchmark designed to measure correctness on fact-seeking questions.
DeepSeek V4 Pro by DeepSeek currently leads with 45%.
2 models in the LuminaBench cohort have a qualifying score on this benchmark.
Yes. This family is ranking-weighted at 3.5% of its capability category. DeepSeek V4 Pro is overall #55.
Related