Seunghyun Lee, David Brumley, Carnegie Mellon University
Category
Knowledge
Version
2026
Direction
higher is better
Ranking use
Reference
Contamination risk
Unknown
A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.