Measuring Short-Form Factuality in Large Language Models
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
Reference2024Active9 models
Data verified 27 Jul 2026 · Methodology 1.8.0
Benchmark score on Measuring Short-Form Factuality in Large Language Models
About Measuring Short-Form Factuality in Large Language Models
Definition and scoring
Organisation
Jason Wei, Najoung Kim, Hyung Won Chung, Yu-An Chung, Siddhartha Papay, Yifeng Lu, Hannaneh Hajishirzi, Luke Zettlemoyer
Category
Knowledge
Version
2024
Direction
higher is better
Ranking use
Reference
Contamination risk
Unknown
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
What does Measuring Short-Form Factuality in Large Language Models measure?
A benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately. Focuses on factual correctness rather than reasoning complexity.
Which model scores highest on Measuring Short-Form Factuality in Large Language Models?
DeepSeek V4 Pro (Max) by DeepSeek currently leads with 57.9%.
How many models are evaluated on Measuring Short-Form Factuality in Large Language Models?
9 models in the LuminaBench cohort have a qualifying score on this benchmark.