| Kimi K3 as published by Z.AI kimi-k3-max | Kimi K3 | 76.5% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Opus 4.8 as published by Z.AI claude-opus-4-8-max | Claude Opus 4.8 | 76.2% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GPT-5.6 Sol as published by Z.AI gpt-5-6-sol-max | GPT-5.6 Sol | 74.9% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Fable 5 w/ fallback as published by Z.AI claude-fable-5-max | Claude Fable 5 | 74.7% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| DeepSeek V4 Pro 0813 as published by Z.AI deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 74.1% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 73% | VerifiedZ.AI GLM-5.3 official comparison table | Score input provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Qwen3.8 Max as published by Z.AI qwen-3-8-max-xhigh | Qwen3.8 Max | 72.5% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.2 as published by Z.AI glm-5-2-max | GLM-5.2 | 59.9% | VerifiedZ.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 75.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 59.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 58% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 56.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 55.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 54.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 53.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 53.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 51.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 50% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 49.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 49% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 48.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 47.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 42.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 40.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 39.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 38% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 36.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 35.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 27.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 26.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 75.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 59.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 58% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 56.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 55.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 54.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 53.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 53.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 51.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 50% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 49.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 49% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 48.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 47.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 42.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 40.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 39.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 38% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 36.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 35.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 27.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 26.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 75.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 59.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 58% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 56.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 55.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 54.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 53.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 53.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 51.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 50% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 49.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 49% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 48.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 47.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 43.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 42.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 40.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 39.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 38% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 36.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 35.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 27.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 26.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Toolathlon-Verified; max reasoning; as published in the Kimi K3 launch blog. kimi-k3-max | Kimi K3 | 73.2% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark 1.1 | 75.6% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.8 | 59.9% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Sol | 58% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 56.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 55.6% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 54.6% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Luna | 53.4% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Terra | 53.1% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5.2 | 48.2% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 46.3% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | MiniMax M2.7 | 46.3% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 43.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 40.7% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 39.8% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5 | 38% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 nano | 35.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 27.8% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Muse Spark 1.1; provider evaluation configuration as published in the Meta Muse Spark 1.1 evaluation report (Toolathlon). muse-spark-1-1-xhigh | Muse Spark 1.1 | 75.6% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Meta AI Muse Spark 1.1 evaluation report ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.5 Flash | 56.5% | May 2026Provider evaluation harness | Score input provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | GPT-5.5 | 55.6% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|