| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run) claude-opus-4-7-max | Claude Opus 4.7 | 42.3% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GLM-5.2 (max) (Artificial Analysis independent run) glm-5-2-max | GLM-5.2 | 41.1% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run) claude-opus-4-6-max | Claude Opus 4.6 | 39.9% | text-only currentArtificial Analysis text-only HLE evaluation | Public reference · not scored independently-verified | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.8-Flash-Next (Artificial Analysis independent run) qwen-3-8-flash-next-aa-unspecified | Qwen3.8-Flash-Next | 38.0% | text-only currentArtificial Analysis text-only HLE evaluation | Public reference · not scored independently-verified | Qwen3.8-Flash-Next individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| LongCat 2.0 (Artificial Analysis independent run) longcat-2-0-default | LongCat-2.0 | 33.7% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Kimi K2.5 (Reasoning) (Artificial Analysis completed independent run) kimi-k2-5-thinking | Kimi K2.5 | 30.7% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Kimi K2.5 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 397B A17B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-397b-thinking | Qwen3.5 397B A17B | 29.0% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Qwen3.5 397B A17B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 122B A10B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-122b-a10b-aa-reasoning-default | Qwen3.5-122B-A10B | 25.2% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Qwen3.5 122B A10B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.6 35B A3B (Reasoning) (Artificial Analysis completed independent run) qwen3-6-35b-a3b-aa-reasoning-default | Qwen3.6-35B-A3B | 22.2% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Qwen3.6 35B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run) gemma-4-26b-a4b-aa-reasoning-default | Gemma 4 26B A4B | 19.3% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Trinity Large Thinking (Artificial Analysis completed independent run) trinity-large-thinking-aa-reasoning-default | Trinity-Large-Thinking | 15.8% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Trinity Large Thinking current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run) deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 14.3% | text-only currentArtificial Analysis text-only HLE evaluation | Public reference · not scored independently-verified | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Medium 3.5 (Artificial Analysis completed independent run) mistral-medium-3-5-128b-aa-reasoning-default | Mistral Medium 3.5 128B | 13.8% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) (Artificial Analysis completed independent run) nemotron-3-nano-30b-aa-reasoning-default | Nemotron 3 Nano 30B | 11.4% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Small 4 (Reasoning) (Artificial Analysis completed independent run) mistral-small-4-reasoning | Mistral Small 4 | 9.9% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Mistral Small 4 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run) deepseek-v3-1-non-reasoning | DeepSeek V3.1 | 6.7% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 mini (Artificial Analysis completed independent run) gpt-4-1-mini-epoch-gpt-4-1-mini-2025-04-14 | GPT-4.1 mini | 5.0% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | GPT-4.1 mini current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Maverick (Artificial Analysis completed independent run) llama-4-maverick-aa-non-reasoning-default | Llama 4 Maverick | 4.9% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Llama 4 Maverick current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 3 27B Instruct (Artificial Analysis completed independent run) gemma-3-27b-aa-non-reasoning-default | Gemma 3 27B | 4.4% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Gemma 3 27B Instruct current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Scout (Artificial Analysis completed independent run) llama-4-scout-aa-non-reasoning-default | Llama 4 Scout | 3.8% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 nano (Artificial Analysis completed independent run) gpt-4-1-nano-epoch-gpt-4-1-nano-2025-04-14 | GPT-4.1 nano | 3.8% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | GPT-4.1 nano current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Fable 5 (max; Artificial Analysis independent run) claude-fable-5-max | Claude Fable 5 | 55.5% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Claude Fable 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Opus 5 (max; Artificial Analysis independent run) claude-opus-5-max | Claude Opus 5 | 54.9% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Claude Opus 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Sol (max; Artificial Analysis independent run) gpt-5-6-sol-max | GPT-5.6 Sol | 49.5% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GPT-5.6 Sol individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Opus 4.8 (max; Artificial Analysis independent run) claude-opus-4-8-max | Claude Opus 4.8 | 48.7% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Claude Opus 4.8 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Kimi K3 (max; Artificial Analysis independent run) kimi-k3-max | Kimi K3 | 46.9% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Kimi K3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.1 (xhigh; Artificial Analysis independent run) muse-spark-1-1-xhigh | Muse Spark 1.1 | 46.2% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.5 (xhigh; Artificial Analysis independent run) gpt-5-5-xhigh | GPT-5.5 | 45.8% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GPT-5.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.2 (xhigh; Artificial Analysis independent run) muse-spark-1-2-xhigh | Muse Spark 1.2 | 45.5% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | Muse Spark 1.2 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.4 (xhigh; Artificial Analysis independent run) gpt-5-4-xhigh | GPT-5.4 | 43.7% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GPT-5.4 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Terra (max; Artificial Analysis independent run) gpt-5-6-terra-max | GPT-5.6 Terra | 42.9% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GPT-5.6 Terra individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Grok 4.6 (high; Artificial Analysis independent run) grok-4-6-high | Grok 4.6 | 42.9% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Grok 4.6 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Grok 4.5 (high; Artificial Analysis independent run) grok-4-5-aa-2-high | Grok 4.5 | 42.7% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Grok 4.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3 (max; Artificial Analysis independent run) glm-5-3-max | GLM-5.3 | 42.3% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GLM-5.3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Sonnet 5 (max; Artificial Analysis independent run) claude-sonnet-5-max | Claude Sonnet 5 | 41.3% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Claude Sonnet 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Pro 0813 (max; Artificial Analysis independent run) deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 41.0% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Public reference · not scored independently-verified | DeepSeek V4 Pro 0813 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.6 Flash (high; Artificial Analysis independent run) gemini-3-6-flash-high | Gemini 3.6 Flash | 40.8% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Gemini 3.6 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3-Flash (max; Artificial Analysis independent run) glm-5-3-flash-max | GLM-5.3-Flash | 39.9% | text-only currentArtificial Analysis text-only HLE evaluation | Score input independently-verified | GLM-5.3-Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Luna (max; Artificial Analysis independent run) gpt-5-6-luna-max | GPT-5.6 Luna | 39.5% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | GPT-5.6 Luna individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.7 Flash (medium; Artificial Analysis independent run) gemini-3-7-flash-medium | Gemini 3.7 Flash | 39.0% | text-only currentArtificial Analysis text-only Humanity's Last Exam evaluation | Score input independently-verified | Gemini 3.7 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run) deepseek-v4-flash-vision-exp-max-harness | DeepSeek V4 Flash Vision Exp | 34.5% | text-only currentArtificial Analysis text-only HLE evaluation | Public reference · not scored independently-verified | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3-Flash (max effort, documented default) glm-5-3-flash-max | GLM-5.3-Flash | 55.3% | rollingsystem:glm-5-3-flash-launch:hle-tools:0 | Public reference · not scored source-checked | GLM-5.3-Flash individual evaluations ↗Observed 2026-08-26 · checked 2026-08-27 |
|---|
| Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison table claude-opus-4-6-max | Claude Opus 4.6 | 40% | text-onlysystem:qwen3-8-flash-next-comparison:cell:language:hle:4 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 34.7% | text-onlysystem:qwen3-8-flash-next-comparison:cell:language:hle:2 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 33.8% | text-onlysystem:qwen3-8-flash-next-comparison:cell:language:hle:3 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 30.8% | text-onlysystem:qwen3-8-flash-next-comparison:cell:language:hle:1 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Claude Opus 4.6 Max as published by Qwen claude-opus-4-6-max | Claude Opus 4.6 | 40% | text-onlyQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 34.7% | text-onlyQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 32.3% | without toolsSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 30.8% | text-onlyQwen3.8-27B official text table | Score input provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 28.8% | without toolsSolar Open 2 card English table | Score input provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 24.3% | without toolsSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 24% | text-onlyQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 12.8% | without toolsSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 11.5% | without toolsSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 11.4% | without toolsSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64 muse-glimmer-30b-high | Muse Glimmer 30B | 22% | text-only / no tools / 2,158 questionsArtificial Analysis HLE text-only evaluation | Public reference · not scored provider-reported | Muse Glimmer Evaluation Methodology ↗Observed 2026-08-10 · checked 2026-08-10 |
|---|
| Qwen3.8 Max (xhigh, default reasoning effort) qwen-3-8-max-xhigh | Qwen3.8 Max | 43.6% | CurrentQwen provider evaluation without tools | Public reference · not scored provider-reported | Qwen3.8 Max official release ↗Observed 2026-08-03 · checked 2026-08-05 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 64.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 64.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 62.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-4-pro-default-medium | GPT-5.4 | 58.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 57.9% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 57.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-5-pro-default-high | GPT-5.5 | 57.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Kimi K3 public default configuration (reasoning_effort=max; temperature=1.0) kimi-k3-max | Kimi K3 | 56% | HLE-Full with general toolsMoonshot Kimi K3 model-card evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-29 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 53.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 53% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 52.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 52.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 52.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 49% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 48% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling-Small; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling-Small | 47.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Agents-A1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Agents-A1 | 47.6% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 47.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 46% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 41.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 41.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 41.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 40.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 40.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 39.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 38.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 37.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 37.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 36.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 35.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 35% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 34.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 34.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 33.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 32.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 31.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 31.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 30.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 30.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 29.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 29.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 28.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 28.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 28.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 28.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 27.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 26.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 26.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 25.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 25.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 24.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 24% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 23.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 23.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 22.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 21.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 20% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 19.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 18.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 18.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 17.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 17.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 16.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 15.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 14.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 14.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 13.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 13% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 12.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 11.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 11.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 11.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 10.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 10.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 10.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 9.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 8.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 8.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 8.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 7.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 7.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 7.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 6.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 6.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 6.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 6.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 6.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 5.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 5.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 5.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 5.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 5.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 4.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 4.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 3.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 3.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 3.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4 Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4 Turbo | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 64.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 64.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 62.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-4-pro-default-medium | GPT-5.4 | 58.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 57.9% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 57.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-5-pro-default-high | GPT-5.5 | 57.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 56% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 53.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 53% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 52.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 52.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 52.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 49% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 48% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Agents-A1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Agents-A1 | 47.6% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 47.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 46% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 41.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 41.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 41.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 40.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 40.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 39.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 38.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 37.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 37.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 36.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 35.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 35% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 34.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 34.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 33.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 32.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 31.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 31.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 30.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 30.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 29.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 29.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 28.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 28.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 28.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 28.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 27.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 26.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 26.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 25.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 25.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 24.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 24% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 23.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 23.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 22.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 21.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 20% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 19.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 18.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 18.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 17.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 17.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 16.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 15.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 14.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 14.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 13.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 13% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 12.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 11.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 11.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 11.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 10.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 10.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 10.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 9.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 8.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 8.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 8.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 7.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 7.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 7.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 6.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 6.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 6.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 6.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 6.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 5.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 5.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 5.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 5.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 5.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 4.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 4.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 3.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 3.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 3.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4 Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4 Turbo | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Humanity's Last Exam with tools as published in the Anthropic Claude Opus 5 launch table. source label without registered configuration ID | Claude Opus 5 | 64.7% | CurrentAnthropic provider evaluation | Score input provider-reported | Introducing Claude Opus 5 ↗Observed 2026-07-24 · checked 2026-07-24 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 64.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 62.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-4-pro-default-medium | GPT-5.4 | 58.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 57.9% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 57.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5 Pro; bulk export does not retain a complete upstream harness configuration. gpt-5-5-pro-default-high | GPT-5.5 | 57.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 56% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 54.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 53.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 53% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 52.3% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 52.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 52.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 50.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 49% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 48% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Agents-A1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Agents-A1 | 47.6% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 47.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 46% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 41.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 41.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 41.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 40.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 40.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 39.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 38.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis Humanity's Last Exam independent evaluation. source label without registered configuration ID | Gemini 3.6 Flash | 38.3% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Gemini 3.6 Flash (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 37.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 37.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 37.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 36.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 35.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 35% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 34.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 34.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 34.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 33.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 32.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 31.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 31.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 30.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 30.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 29.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 29.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 28.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 28.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 28.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 28.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 27.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 26.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 26.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 26.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 25.5% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 25.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 24.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 24% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 23.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 23.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 23.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 22.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 21.4% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 20% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 19.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 19.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 18.8% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 18.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 17.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis Humanity's Last Exam independent evaluation. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 17.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Gemini 3.5 Flash-Lite (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 17.2% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 17% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 16.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 15.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 14.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 14.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 14.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 14.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 13.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 13% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 12.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 11.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 11.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 11.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 10.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 10.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 10.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 9.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 9.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 8.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 8.1% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 8.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 7.7% | 2025Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 7.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 7.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 6.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 6.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 6.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 6.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 6.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 5.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 5.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 5.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 5.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 5.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 5.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 4.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 4.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 4.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 4.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 4.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 4.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 4.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 3.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 3.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 3.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 3.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 3.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4 Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4 Turbo | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 3.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 3.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| HLE-Full with tools; max reasoning; provider-published Kimi K3 evaluation. kimi-k3-max | Kimi K3 | 56% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Artificial Analysis HLE evaluation; configuration Adaptive Reasoning, Max Effort, Opus 4.8 Fallback. claude-fable-5-max | Claude Fable 5 | 53.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Humanity's Last Exam leaderboard (AA) ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| Artificial Analysis HLE evaluation; configuration max. gpt-5-6-sol-max | GPT-5.6 Sol | 47.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Humanity's Last Exam leaderboard (AA) ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| Artificial Analysis HLE evaluation; configuration Adaptive Reasoning, Max Effort. claude-opus-4-8-max | Claude Opus 4.8 | 45.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Humanity's Last Exam leaderboard (AA) ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| Muse Spark 1.1 (xhigh); Artificial Analysis independent evaluation configuration as published. muse-spark-1-1-xhigh | Muse Spark 1.1 | 45.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Humanity's Last Exam leaderboard (AA) ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Mythos 5 | 64.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark 1.1 | 62.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.8 | 57.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 5 | 57.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. gpt-5-5-pro-default-high | GPT-5.5 | 57.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5.2 | 54.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Fable 5 | 53.3% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-fable-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.6 | 53% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5.1 | 52.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 52.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 52.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5 | 50.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark | 50.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 4.6 | 49% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | MiMo-V2.5-Pro | 48% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Sol | 47.2% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-sol ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Opus 4.8 | 45.7% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-opus-4-8 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 44.7% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gemini-3-1-pro-preview ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.5 | 44.3% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Terra | 41.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-terra ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.4 | 41.6% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Max | 41.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.5 | 40.3% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 40.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.3-Codex | 39.9% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-3-codex ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.5 Flash | 39.9% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gemini-3-5-flash-medium ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Sonnet 5 | 39.6% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-sonnet-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Qwen3.7-Max | 38.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for qwen3-7-max ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 nano | 37.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Luna | 37.2% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-luna ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | MiniMax M3 | 37.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for minimax-m3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 35.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Pro | 35.9% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-pro ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 35.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.3 | 35.0% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Grok 4.3 | 35% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 34.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 Codex (xhigh) gpt-5-2-codex-xhigh | GPT-5.2 Codex | 33.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.7 Code source label without registered configuration ID | Kimi K2.7 Code | 32.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.20 0309 v2 (Reasoning) source label without registered configuration ID | Grok 4.20 | 32.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Flash | 32.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-flash ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 30.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 30.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 Max Preview source label without registered configuration ID | Qwen3.6 Max | 28.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 28.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 27.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 26.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nemotron 3 Ultra 550B A55B (Reasoning) source label without registered configuration ID | Nemotron 3 Ultra | 26.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 26.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemma 4 31B | 26.5% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3 Max Thinking source label without registered configuration ID | Qwen3 Max | 26.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 25.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5 source label without registered configuration ID | MiMo-V2.5 | 25.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 25.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 23.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 122B A10B (Reasoning) source label without registered configuration ID | Qwen3.5 122B | 23.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 Thinking source label without registered configuration ID | Kimi K2 Thinking | 22.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 22.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 27B (Reasoning) source label without registered configuration ID | Qwen3.5 27B | 22.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.1 source label without registered configuration ID | MiniMax M2.1 | 22.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 27B (Reasoning) source label without registered configuration ID | Qwen3.6 27B | 21.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 21.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 20.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 35B A3B (Reasoning) source label without registered configuration ID | Qwen3.5 35B | 19.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| NVIDIA Nemotron 3 Super 120B A12B (Reasoning) source label without registered configuration ID | Nemotron 3 Super | 19.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.5 source label without registered configuration ID | MiniMax M2.5 | 19.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 26B A4B (Reasoning) source label without registered configuration ID | Gemma 4 26B | 18.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 17.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 17.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 17.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 15.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 12B (Reasoning) source label without registered configuration ID | Gemma 4 12B Unified | 14.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 14.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 Omni Plus source label without registered configuration ID | Qwen3.5 Omni Plus | 13.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 13.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Medium 3.5 | 12.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 12.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2 source label without registered configuration ID | MiniMax M2 | 12.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 11.9% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4 | 11.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Command A+ source label without registered configuration ID | Command A+ | 11.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 3 mini Reasoning (high) source label without registered configuration ID | Grok 3 mini | 11.1% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 3.7 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 3.7 | 10.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 9.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Small 4 | 9.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-small-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 8.9% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Ling-2.6-1T source label without registered configuration ID | Ling 2.6 1T | 8.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 8.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o1 source label without registered configuration ID | o1 | 7.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 7.7% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok Code Fast 1 source label without registered configuration ID | Grok Code Fast 1 | 7.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 0905 source label without registered configuration ID | Kimi K2 0905 | 6.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Haiku 4.5 | 4.3% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for claude-4-5-haiku ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Large 3 | 4.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-large-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Muse Spark 1.1; provider evaluation configuration as published in the Meta Muse Spark 1.1 evaluation report (Humanity's Last Exam (with tools as published)). muse-spark-1-1-xhigh | Muse Spark 1.1 | 62.1% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Meta AI Muse Spark 1.1 evaluation report ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|
| tencent/Hy3 instruct weights; configuration as recorded on the Hugging Face evaluation results entry. source label without registered configuration ID | Hy3 | 53.2% | CurrentProvider / community evaluation result attached on Hugging Face | Score input provider-reported | Tencent Hy3 model card on Hugging Face ↗Observed 2026-07-06 · checked 2026-07-16 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.1 Pro Preview | 44.4% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | GPT-5.5 | 41.4% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.5 Flash | 40.2% | CurrentProvider evaluation harness | Score input provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|