| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 97% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 95.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.5 | 87% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 85.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 82.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 42.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 40.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 97% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 95.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.5 | 87% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 85.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 82.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 42.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 40.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 97% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 96.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 95.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.5 | 87% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 85.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 82.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 42.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 40.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 99% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 98.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 96.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 95% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 Thinking source label without registered configuration ID | Kimi K2 Thinking | 94.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 94.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 92.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 92% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 90.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 89.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 89.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 89% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 88.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 88% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 87.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 86% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 85% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 3 mini Reasoning (high) source label without registered configuration ID | Grok 3 mini | 84.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.1 source label without registered configuration ID | MiniMax M2.1 | 82.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 80.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2 source label without registered configuration ID | MiniMax M2 | 78.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 78.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 74.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4 | 73.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 0905 source label without registered configuration ID | Kimi K2 0905 | 57.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 3.7 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 3.7 | 56.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok Code Fast 1 source label without registered configuration ID | Grok Code Fast 1 | 43.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|