Artificial Analysis private agentic knowledge-work evaluation with Elo over rubric pass rate, analytical quality, and presentation quality on business deliverables.
Artificial Analysis private agentic knowledge-work evaluation with Elo over rubric pass rate, analytical quality, and presentation quality on business deliverables. Results stay tied to the exact model variant and evaluation system. Multiple systems for the same model use the best published score on this page; overall Lumina scoring uses the median of ranking-eligible rows.
Artificial Analysis private agentic knowledge-work evaluation with Elo over rubric pass rate, analytical quality, and presentation quality on business deliverables.
Which model scores highest on AA-Briefcase?
Kimi K3 by Moonshot AI currently leads with 152.7 elo-proxy.
How many models are evaluated on AA-Briefcase?
4 models in the LuminaBench cohort have a qualifying score on this benchmark.
Does this affect overall Lumina rank?
No. This benchmark is display-only and does not enter the overall Lumina composite.