| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 86.2% | standardSolar Open 2 card English table | Score input provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 85.9% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 84.6% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 81.2% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 80.4% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 79% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95 nemotron-3-5-lightning-30b-a3b-default | NVIDIA Nemotron 3.5 Lightning 30B-A3B | 81.9% | CurrentNVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | Score input provider-reported | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model card ↗Observed 2026-08-11 · checked 2026-08-12 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.5 397B A17B | 87.8% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Nemotron 3 Ultra | 86.8% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.5 122B | 86.7% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.6 27B | 86.2% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.5 27B | 86.1% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.5 35B | 85.3% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | GLM-4.7 | 84.3% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Gemma 4 26B | 82.6% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public mmlu-pro leaderboard page (verified 2026-07-21). source label without registered configuration ID | Gemma 4 12B Unified | 77.2% | CurrentBenchLM aggregated public evaluation | Score input source-checked | MMLU-Pro Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Max | 89.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 89.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 88.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 88.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 87.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5 | 85.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemma 4 31B | 85.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 83% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 82.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.6 | 82% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 4.6 | 79.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|