| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 16.6% | standardSolar Open 2 card English table | Score input provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 13.4% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 13.2% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 6.1% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 2.4% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 1.6% | standardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Fable 5 Max claude-fable-5-max | Claude Fable 5 | 59.2% | CurrentxAI Grok 4.6 official release comparison table | Relative comparison · not standard direct provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Grok 4.6 High grok-4-6-high | Grok 4.6 | 57.5% | CurrentxAI Grok 4.6 official release comparison table | Score input provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| GPT-5.6 Sol Max gpt-5-6-sol-max | GPT-5.6 Sol | 56.7% | CurrentxAI Grok 4.6 official release comparison table | Relative comparison · not standard direct provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Grok 4.5 High grok-4-5-aa-2-high | Grok 4.5 | 47.1% | CurrentxAI Grok 4.6 official release comparison table | Relative comparison · not standard direct provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 47.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 41.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 38.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 37.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 35.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 33% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 28.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 28.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 24.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 24.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 22.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 14.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 12.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 10.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 2.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 0.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 47.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 41.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 38.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 37.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 35.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 33% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 28.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 28.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 24.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 24.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 22.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 14.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 12.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 10.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 2.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 0.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 47.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 41.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 38.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 37.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 35.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 33% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 28.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 28.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 24.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 24.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 22.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 15.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 14.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 12.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 11.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 10.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 2.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 0.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| APEX-Agents; max reasoning; provider-published Kimi K3 evaluation. kimi-k3-max | Kimi K3 | 37.6% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Gemini 3.5 Flash (high) gemini-3-5-flash-high | Gemini 3.5 Flash | 47.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.5 (xhigh) gpt-5-5-xhigh | GPT-5.5 | 37.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5.2 (max) glm-5-2-max | GLM-5.2 | 33.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (xhigh) gpt-5-4-xhigh | GPT-5.4 | 33.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Opus 4.6 (Adaptive Reasoning, Max Effort) claude-opus-4-6-max | Claude Opus 4.6 | 33.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.1 Pro Preview source label without registered configuration ID | Gemini 3.1 Pro Preview | 32.0% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 28.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 28.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) claude-sonnet-4-6-max | Claude Sonnet 4.6 | 28.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3 Flash Preview (Reasoning) source label without registered configuration ID | Gemini 3 Flash | 27.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 nano (xhigh) gpt-5-4-nano-xhigh | GPT-5.4 nano | 24.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V4 Pro (Reasoning, Max Effort) deepseek-v4-pro-max | DeepSeek V4 Pro | 24.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.7 Plus source label without registered configuration ID | Qwen3.7-Plus | 22.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.3 (high) source label without registered configuration ID | Grok 4.3 | 17.0% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 15.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 14.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5 (Reasoning) source label without registered configuration ID | GLM-5 | 14.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.1 Flash-Lite source label without registered configuration ID | Gemini 3.1 Flash-Lite | 12.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.5 (Reasoning) source label without registered configuration ID | Kimi K2.5 | 11.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.7 source label without registered configuration ID | MiniMax M2.7 | 10.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5-Pro source label without registered configuration ID | MiMo-V2.5-Pro | 2.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| NVIDIA Nemotron 3 Super 120B A12B (Reasoning) source label without registered configuration ID | Nemotron 3 Super | 1.8% | CurrentArtificial Analysis independent evaluation | Score input source-checked | APEX-Agents-AA Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| tencent/Hy3 instruct weights; configuration as recorded on the Hugging Face evaluation results entry. source label without registered configuration ID | Hy3 | 25.6% | CurrentProvider / community evaluation result attached on Hugging Face | Score input provider-reported | Tencent Hy3 model card on Hugging Face ↗Observed 2026-07-06 · checked 2026-07-16 |
|---|