| Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64 muse-glimmer-30b-high | Muse Glimmer 30B | 74% | 1,730-question setArtificial Analysis MMMU-Pro documented evaluation | Public reference · not scored provider-reported | Muse Glimmer Evaluation Methodology ↗Observed 2026-08-10 · checked 2026-08-10 |
|---|
| Qwen3.8 Max (xhigh, default reasoning effort) qwen-3-8-max-xhigh | Qwen3.8 Max | 82.3% | CurrentQwen provider evaluation | Public reference · not scored provider-reported | Qwen3.8 Max official release ↗Observed 2026-08-03 · checked 2026-08-05 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 83.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 83.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 80.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 80.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 79.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 79% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 78.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 78.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 78.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 78.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 78% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 76.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 75.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 74.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 74.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 74.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 74.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 74% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 73.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 71.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 70.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 70.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 69.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 69.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 69.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 67.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 65.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 64.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 62.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 62.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 61.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 61.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 55% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 53.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 53% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 52.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 51.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 48.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 48% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 44.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 44.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 41.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 40.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 30.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 83.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 83.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 80.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 80.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 79.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 79% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 78.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 78.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 78.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 78.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 78% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 76.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 75.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 74.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 74.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 74.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 74.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 74% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 73.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 71.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 70.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 70.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 69.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 69.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 69.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 67.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 65.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 64.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 62.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 62.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 61.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 61.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 55% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 53.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 53% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 52.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 51.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 48.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 48% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 44.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 44.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 41.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 40.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 30.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 83.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 83.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis MMMU-Pro independent evaluation. source label without registered configuration ID | Gemini 3.6 Flash | 83.2% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Gemini 3.6 Flash (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 80.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 80.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 80.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 79.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 79% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis MMMU-Pro independent evaluation. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 79% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Gemini 3.5 Flash-Lite (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 78.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 78.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 78.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 78.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 78.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 78% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 76.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 75.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 75.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 75.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 75% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 74.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 74.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 74.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 74.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 74% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 73.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 72.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 71.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 70.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 70.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 69.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 69.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 69.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 67.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 65.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 64.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 62.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 62.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 61.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 61.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 56.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 55% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 53.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 53% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 52.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 51.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 48.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 48% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 44.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 44.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 41.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 40.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 30.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| MMMU-Pro official protocol; original input order; images prepended to text; max reasoning; average of three runs. kimi-k3-max | Kimi K3 | 81.6% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 83.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.5 Flash | 83.9% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gemini-3-5-flash-medium ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 83.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Sol | 83.4% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-6-sol ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Sol | 83% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 82.4% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gemini-3-1-pro-preview ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 81.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 81.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Terra | 80.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Terra | 80.7% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-6-terra ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.5 | 80.4% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for grok-4-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Grok 4.5 | 80.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark | 80.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.5 | 79.9% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 79.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 79% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 78.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Luna | 78.5% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-6-luna ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | MiniMax M3 | 78.5% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for minimax-m3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.3-Codex | 78.5% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-3-codex ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 78.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.3-Codex | 78.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.4 | 78.4% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for gpt-5-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Luna | 78.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | MiniMax M3 | 78.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Grok 4.3 | 78.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.3 | 78.1% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for grok-4-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 5 | 77.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.6 | 77.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 77.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Sonnet 5 | 77.3% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for claude-sonnet-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemma 4 31B | 76.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 Codex (xhigh) gpt-5-2-codex-xhigh | GPT-5.2 Codex | 76.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 75.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5 source label without registered configuration ID | MiMo-V2.5 | 75.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 27B (Reasoning) source label without registered configuration ID | Qwen3.5 27B | 75.0% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 122B A10B (Reasoning) source label without registered configuration ID | Qwen3.5 122B | 75.0% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 74.9% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 27B (Reasoning) source label without registered configuration ID | Qwen3.6 27B | 74.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.20 0309 v2 (Reasoning) source label without registered configuration ID | Grok 4.20 | 74.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 74.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 73.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 73.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 73.1% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 35B A3B (Reasoning) source label without registered configuration ID | Qwen3.5 35B | 72.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 70.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 Omni Plus source label without registered configuration ID | Qwen3.5 Omni Plus | 70.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 70.1% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 12B (Reasoning) source label without registered configuration ID | Gemma 4 12B Unified | 69.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 26B A4B (Reasoning) source label without registered configuration ID | Gemma 4 26B | 69.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 69.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 68.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 68.8% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 68.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 67.9% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 nano | 66.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Medium 3.5 | 64.9% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 64.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Command A+ source label without registered configuration ID | Command A+ | 63.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 61.8% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 61.8% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Small 4 | 56.8% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for mistral-small-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Mistral Small 4 | 56.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Mistral Large 3 | 55.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Large 3 | 55.7% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for mistral-large-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Haiku 4.5 | 55.1% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for claude-4-5-haiku ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.5 Flash | 83.6% | CurrentProvider evaluation harness | Score input provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | GPT-5.5 | 81.2% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.1 Pro Preview | 80.5% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| GPT-4o mini source label without registered configuration ID | GPT-4o mini | 37.6% | 2024MMMU-Pro official owner evaluation (paper / leaderboard_data.json unlabeled row) | Score input official-leaderboard | MMMU permanent refresh source ↗Observed 2024-07-18 · checked 2026-08-18 |
|---|
| Gemini 1.5 Pro (0523) source label without registered configuration ID | Gemini 1.5 Pro (May '24) | 43.5% | 2024MMMU-Pro official owner evaluation (paper / leaderboard_data.json unlabeled row) | Score input official-leaderboard | MMMU permanent refresh source ↗Observed 2024-05-23 · checked 2026-08-18 |
|---|
| GPT-4o (0513) source label without registered configuration ID | GPT-4o (May '24) | 51.9% | 2024MMMU-Pro official owner evaluation (paper / leaderboard_data.json unlabeled row) | Score input official-leaderboard | MMMU permanent refresh source ↗Observed 2024-05-13 · checked 2026-08-18 |
|---|