Navigate the site or open source-checked model and benchmark registry records.
Navigate
Lumina Bench
News
Research
Independent rankings. Commercial relationships never affect scoring.
Anthropic · stable
Anthropic's most capable widely released model, documented for long-running agent workloads.
#2 · 83.0 score · Confidence B