| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 41.7% | HardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 34.1% | HardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 33.3% | HardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 28.3% | HardSolar Open 2 card English table | Public reference · not scored provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 25% | HardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 2.3% | HardSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 65.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 62.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 58.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 57.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 46.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 43.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 35.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 25% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 23.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 65.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 62.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 60.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 58.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 57.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 46.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 43.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 35.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 25% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 65.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 62.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 60.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 58.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 57.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 46.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 43.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 43.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 40.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 36.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 35.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 33.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 25% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Kimi K3; Artificial Analysis Terminal-Bench Hard / Terminal-Bench v2.1 independent evaluation. source label without registered configuration ID | Kimi K3 | 85% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Kimi K3 Intelligence, Performance & Price Analysis ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| GPT-5.6 Sol (max) gpt-5-6-sol-max | GPT-5.6 Sol | 65.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) claude-fable-5-max | Claude Fable 5 | 62.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.5 (xhigh) gpt-5-5-xhigh | GPT-5.5 | 60.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Opus 4.8 (Adaptive Reasoning, Max Effort) claude-opus-4-8-max | Claude Opus 4.8 | 58.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (xhigh) gpt-5-4-xhigh | GPT-5.4 | 57.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.6 Terra (max) gpt-5-6-terra-max | GPT-5.6 Terra | 57.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.1 Pro Preview source label without registered configuration ID | Gemini 3.1 Pro Preview | 53.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) claude-sonnet-4-6-max | Claude Sonnet 4.6 | 53.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.3 Codex (xhigh) gpt-5-3-codex-xhigh | GPT-5.3-Codex | 53.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 52.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) claude-opus-4-7-max | Claude Opus 4.7 | 51.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.7 Max source label without registered configuration ID | Qwen3.7-Max | 50.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5.2 (max) glm-5-2-max | GLM-5.2 | 50.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.7 Plus source label without registered configuration ID | Qwen3.7-Plus | 47.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 47.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Opus 4.5 (Reasoning) source label without registered configuration ID | Claude Opus 4.5 | 47.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude Opus 4.6 (Adaptive Reasoning, Max Effort) claude-opus-4-6-max | Claude Opus 4.6 | 46.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V4 Pro (Reasoning, Max Effort) deepseek-v4-pro-max | DeepSeek V4 Pro | 46.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.1 (high) gpt-5-1-high | GPT-5.1 | 45.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Muse Spark source label without registered configuration ID | Muse Spark | 45.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.7 Code source label without registered configuration ID | Kimi K2.7 Code | 44.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 Plus source label without registered configuration ID | Qwen3.6 Plus | 43.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 43.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 Max Preview source label without registered configuration ID | Qwen3.6 Max | 43.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5 (Reasoning) source label without registered configuration ID | GLM-5 | 43.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5.1 (Reasoning) source label without registered configuration ID | GLM-5.1 | 43.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5-Pro source label without registered configuration ID | MiMo-V2.5-Pro | 43.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M3 source label without registered configuration ID | MiniMax M3 | 42.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 nano (xhigh) gpt-5-4-nano-xhigh | GPT-5.4 nano | 42.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5 source label without registered configuration ID | MiMo-V2.5 | 41.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2-Pro source label without registered configuration ID | MiMo-V2-Pro | 40.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 40.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.5 Flash (high) gemini-3-5-flash-high | Gemini 3.5 Flash | 40.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.7 source label without registered configuration ID | MiniMax M2.7 | 39.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3 Flash Preview (Reasoning) source label without registered configuration ID | Gemini 3 Flash | 38.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.3 (high) source label without registered configuration ID | Grok 4.3 | 37.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 37.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.20 0309 v2 (Reasoning) source label without registered configuration ID | Grok 4.20 | 37.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 37.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 Codex (xhigh) gpt-5-2-codex-xhigh | GPT-5.2 Codex | 37.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 37.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 31B (Reasoning) source label without registered configuration ID | Gemma 4 31B | 36.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nemotron 3 Ultra 550B A55B (Reasoning) source label without registered configuration ID | Nemotron 3 Ultra | 36.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V4 Flash (Reasoning, Max Effort) deepseek-v4-flash-max | DeepSeek V4 Flash | 35.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 35.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 35.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2-Omni source label without registered configuration ID | MiMo-V2-Omni | 34.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.5 (Reasoning) source label without registered configuration ID | Kimi K2.5 | 34.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 27B (Reasoning) source label without registered configuration ID | Qwen3.6 27B | 34.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.5 source label without registered configuration ID | MiniMax M2.5 | 34.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 34.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Mistral Medium 3.5 source label without registered configuration ID | Mistral Medium 3.5 | 33.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-5-Turbo source label without registered configuration ID | GLM-5-Turbo | 33.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 32.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 27B (Reasoning) source label without registered configuration ID | Qwen3.5 27B | 32.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM 5V Turbo (Reasoning) source label without registered configuration ID | GLM-5V-Turbo | 32.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 31.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4 | 31.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 122B A10B (Reasoning) source label without registered configuration ID | Qwen3.5 122B | 31.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 Thinking source label without registered configuration ID | Kimi K2 Thinking | 31.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 31.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Ling-2.6-1T source label without registered configuration ID | Ling 2.6 1T | 31.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 30.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 28.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.1 source label without registered configuration ID | MiniMax M2.1 | 28.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| NVIDIA Nemotron 3 Super 120B A12B (Reasoning) source label without registered configuration ID | Nemotron 3 Super | 28.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Haiku (Reasoning) source label without registered configuration ID | Claude Haiku 4.5 | 27.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 35B A3B (Reasoning) source label without registered configuration ID | Qwen3.5 35B | 26.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 26.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2 source label without registered configuration ID | MiniMax M2 | 25.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Command A+ source label without registered configuration ID | Command A+ | 25% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 25% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.1 Flash-Lite source label without registered configuration ID | Gemini 3.1 Flash-Lite | 24.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 24.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3 Max Thinking source label without registered configuration ID | Qwen3 Max | 24.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 0905 source label without registered configuration ID | Kimi K2 0905 | 23.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 Omni Plus source label without registered configuration ID | Qwen3.5 Omni Plus | 21.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 3.7 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 3.7 | 21.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 18.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 12B (Reasoning) source label without registered configuration ID | Gemma 4 12B Unified | 18.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 3 mini Reasoning (high) source label without registered configuration ID | Grok 3 mini | 17.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok Code Fast 1 source label without registered configuration ID | Grok Code Fast 1 | 17.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Mistral Small 4 (Reasoning) source label without registered configuration ID | Mistral Small 4 | 17.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 16.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Mistral Large 3 source label without registered configuration ID | Mistral Large 3 | 15.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 15.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 26B A4B (Reasoning) source label without registered configuration ID | Gemma 4 26B | 13.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o1 source label without registered configuration ID | o1 | 12.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|