| Qwen3.8-Flash-Next (Artificial Analysis independent run) qwen-3-8-flash-next-aa-unspecified | Qwen3.8-Flash-Next | 92.3% | diamondArtificial Analysis GPQA Diamond evaluation | Public reference · not scored independently-verified | Qwen3.8-Flash-Next individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run) claude-opus-4-7-max | Claude Opus 4.7 | 91.4% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Opus 4.6 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run) claude-opus-4-6-max | Claude Opus 4.6 | 89.6% | diamondArtificial Analysis GPQA Diamond evaluation | Public reference · not scored independently-verified | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GLM-5.2 (max) (Artificial Analysis independent run) glm-5-2-max | GLM-5.2 | 89.5% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 397B A17B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-397b-thinking | Qwen3.5 397B A17B | 89.3% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Qwen3.5 397B A17B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Kimi K2.5 (Reasoning) (Artificial Analysis completed independent run) kimi-k2-5-thinking | Kimi K2.5 | 87.9% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Kimi K2.5 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 122B A10B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-122b-a10b-aa-reasoning-default | Qwen3.5-122B-A10B | 85.7% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Qwen3.5 122B A10B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.6 35B A3B (Reasoning) (Artificial Analysis completed independent run) qwen3-6-35b-a3b-aa-reasoning-default | Qwen3.6-35B-A3B | 84.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Qwen3.6 35B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run) gemma-4-26b-a4b-aa-reasoning-default | Gemma 4 26B A4B | 79.2% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| LongCat 2.0 (Artificial Analysis independent run) longcat-2-0-default | LongCat-2.0 | 78.0% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| DeepSeek V3.1 (Reasoning) (Artificial Analysis independent run) deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 77.9% | diamondArtificial Analysis GPQA Diamond evaluation | Public reference · not scored independently-verified | DeepSeek V3.1 (Reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Small 4 (Reasoning) (Artificial Analysis completed independent run) mistral-small-4-reasoning | Mistral Small 4 | 76.9% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Mistral Small 4 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) (Artificial Analysis completed independent run) nemotron-3-nano-30b-aa-reasoning-default | Nemotron 3 Nano 30B | 75.7% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Trinity Large Thinking (Artificial Analysis completed independent run) trinity-large-thinking-aa-reasoning-default | Trinity-Large-Thinking | 75.2% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Trinity Large Thinking current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Medium 3.5 (Artificial Analysis completed independent run) mistral-medium-3-5-128b-aa-reasoning-default | Mistral Medium 3.5 128B | 74.8% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| DeepSeek V3.1 (Non-reasoning) (Artificial Analysis independent run) deepseek-v3-1-non-reasoning | DeepSeek V3.1 | 73.5% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | DeepSeek V3.1 (Non-reasoning) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Maverick (Artificial Analysis completed independent run) llama-4-maverick-aa-non-reasoning-default | Llama 4 Maverick | 67.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Llama 4 Maverick current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 mini (Artificial Analysis completed independent run) gpt-4-1-mini-epoch-gpt-4-1-mini-2025-04-14 | GPT-4.1 mini | 66.4% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-4.1 mini current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Scout (Artificial Analysis completed independent run) llama-4-scout-aa-non-reasoning-default | Llama 4 Scout | 58.7% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 nano (Artificial Analysis completed independent run) gpt-4-1-nano-epoch-gpt-4-1-nano-2025-04-14 | GPT-4.1 nano | 51.2% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-4.1 nano current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 3 27B Instruct (Artificial Analysis completed independent run) gemma-3-27b-aa-non-reasoning-default | Gemma 3 27B | 42.8% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Gemma 3 27B Instruct current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Grok 4.6 (high; Artificial Analysis independent run) grok-4-6-high | Grok 4.6 | 94.9% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Grok 4.6 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Sol (max; Artificial Analysis independent run) gpt-5-6-sol-max | GPT-5.6 Sol | 94.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-5.6 Sol individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.5 (xhigh; Artificial Analysis independent run) gpt-5-5-xhigh | GPT-5.5 | 93.5% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-5.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Kimi K3 (max; Artificial Analysis independent run) kimi-k3-max | Kimi K3 | 93.5% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Kimi K3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Opus 5 (max; Artificial Analysis independent run) claude-opus-5-max | Claude Opus 5 | 93.2% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Claude Opus 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Grok 4.5 (high; Artificial Analysis independent run) grok-4-5-aa-2-high | Grok 4.5 | 93.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Grok 4.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Pro 0813 (max; Artificial Analysis independent run) deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 92.8% | diamondArtificial Analysis GPQA Diamond evaluation | Public reference · not scored independently-verified | DeepSeek V4 Pro 0813 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.6 Flash (high; Artificial Analysis independent run) gemini-3-6-flash-high | Gemini 3.6 Flash | 92.8% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Gemini 3.6 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Fable 5 (max; Artificial Analysis independent run) claude-fable-5-max | Claude Fable 5 | 92.6% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Claude Fable 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Terra (max; Artificial Analysis independent run) gpt-5-6-terra-max | GPT-5.6 Terra | 92.5% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-5.6 Terra individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.7 Flash (medium; Artificial Analysis independent run) gemini-3-7-flash-medium | Gemini 3.7 Flash | 92.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Gemini 3.7 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Opus 4.8 (max; Artificial Analysis independent run) claude-opus-4-8-max | Claude Opus 4.8 | 92.0% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Claude Opus 4.8 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.4 (xhigh; Artificial Analysis independent run) gpt-5-4-xhigh | GPT-5.4 | 92.0% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-5.4 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3 (max; Artificial Analysis independent run) glm-5-3-max | GLM-5.3 | 91.7% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GLM-5.3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run) deepseek-v4-flash-vision-exp-max-harness | DeepSeek V4 Flash Vision Exp | 91.3% | diamondArtificial Analysis GPQA Diamond evaluation | Public reference · not scored independently-verified | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3-Flash (max; Artificial Analysis independent run) glm-5-3-flash-max | GLM-5.3-Flash | 91.2% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GLM-5.3-Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Sonnet 5 (max; Artificial Analysis independent run) claude-sonnet-5-max | Claude Sonnet 5 | 91.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Claude Sonnet 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Luna (max; Artificial Analysis independent run) gpt-5-6-luna-max | GPT-5.6 Luna | 91.1% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | GPT-5.6 Luna individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.2 (xhigh; Artificial Analysis independent run) muse-spark-1-2-xhigh | Muse Spark 1.2 | 90.4% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Muse Spark 1.2 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.1 (xhigh; Artificial Analysis independent run) muse-spark-1-1-xhigh | Muse Spark 1.1 | 89.8% | diamondArtificial Analysis GPQA Diamond evaluation | Score input independently-verified | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Qwen3.8-Flash-Next (xhigh default thinking configuration) qwen-3-8-flash-next-xhigh | Qwen3.8-Flash-Next | 91.7% | Diamondsystem:qwen3-8-flash-next:gpqa | Score input source-checked | Qwen3.8-Flash-Next current official launch page ↗Observed 2026-08-26 · checked 2026-08-29 |
|---|
| Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison table claude-opus-4-6-max | Claude Opus 4.6 | 91.3% | Diamondsystem:qwen3-8-flash-next-comparison:cell:language:gpqa:4 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 90.8% | Diamondsystem:qwen3-8-flash-next-comparison:cell:language:gpqa:3 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 90.3% | Diamondsystem:qwen3-8-flash-next-comparison:cell:language:gpqa:2 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 89.2% | Diamondsystem:qwen3-8-flash-next-comparison:cell:language:gpqa:1 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 90.9% | 1.0.11Epoch AI benchmark runner | Score input independently-verified | Epoch AI permanent refresh source ↗Observed 2026-08-24 · checked 2026-09-01 |
|---|
| Claude Opus 4.6 Max as published by Qwen claude-opus-4-6-max | Claude Opus 4.6 | 91.3% | DiamondQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 90.3% | DiamondQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 89.2% | DiamondQwen3.8-27B official text table | Score input provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 88.9% | DiamondSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 87.8% | DiamondQwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 86.3% | DiamondSolar Open 2 card English table | Score input provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 83% | DiamondSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 77.5% | DiamondSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 75.6% | DiamondSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 66.2% | DiamondSolar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Gemini 3.7 Flash (high) gemini-3-7-flash-high | Gemini 3.7 Flash | 94.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-14 · checked 2026-08-17 |
|---|
| Grok 4.6 (xhigh) grok-4-6-xhigh | Grok 4.6 | 93.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-14 · checked 2026-08-17 |
|---|
| Inkling Small (xhigh) inkling-small-epoch-inkling-small-xhigh | Inkling-Small | 88.5% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-14 · checked 2026-08-18 |
|---|
| Grok 4.6 (high) grok-4-6-high | Grok 4.6 | 94.0% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-12 · checked 2026-08-17 |
|---|
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95 nemotron-3-5-lightning-30b-a3b-default | NVIDIA Nemotron 3.5 Lightning 30B-A3B | 75.4% | CurrentNVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | Score input provider-reported | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model card ↗Observed 2026-08-11 · checked 2026-08-12 |
|---|
| MiniMax-M3 minimax-m3-epoch-minimax-m3 | MiniMax M3 | 90.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-10 · checked 2026-08-17 |
|---|
| GLM-5.1 glm-5-1-epoch-glm-5-1 | GLM-5.1 | 89.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-10 · checked 2026-08-17 |
|---|
| Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64 muse-glimmer-30b-high | Muse Glimmer 30B | 83.5% | 198-question Diamond setArtificial Analysis GPQA Diamond documented evaluation | Public reference · not scored provider-reported | Muse Glimmer Evaluation Methodology ↗Observed 2026-08-10 · checked 2026-08-10 |
|---|
| GLM-5.2 (no thinking) glm-5-2-epoch-glm-5-2-none | GLM-5.2 | 71.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-10 · checked 2026-08-17 |
|---|
| Kimi K3 (high) kimi-k3-epoch-kimi-k3-high | Kimi K3 | 91.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.7 Max qwen-3-7-max-epoch-qwen3-7-max | Qwen3.7-Max | 90.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.6 Sol (low) gpt-5-6-sol-low | GPT-5.6 Sol | 89.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.7 Plus (thinking) qwen-3-7-plus-epoch-qwen3-7-plus | Qwen3.7-Plus | 87.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Kimi K2.7 Code kimi-k2-7-code-epoch-kimi-k2-7-code | Kimi K2.7 Code | 87.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.6 Terra (low) gpt-5-6-terra-low | GPT-5.6 Terra | 87.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.6 Max Preview (thinking) qwen3-6-max-epoch-qwen3-6-max-preview | Qwen3.6 Max | 87.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 86.9% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.5 397B-A17B (no thinking) qwen3-5-397b-epoch-qwen3-5-397b-a17b-none | Qwen3.5 397B A17B | 86.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Gemini 3.6 Flash (minimal) gemini-3-6-flash-epoch-gemini-3-6-flash-minimal | Gemini 3.6 Flash | 85.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.6 27B (thinking) qwen3-6-27b-epoch-qwen3-6-27b | Qwen3.6 27B | 85.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.5 397B-A17B (thinking) qwen3-5-397b-epoch-qwen3-5-397b-a17b | Qwen3.5 397B A17B | 85.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Kimi K3 (low) kimi-k3-low | Kimi K3 | 84.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.6 27B (no thinking) qwen3-6-27b-epoch-qwen3-6-27b-none | Qwen3.6 27B | 84.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.6 Sol (none) gpt-5-6-sol-epoch-gpt-5-6-sol-none | GPT-5.6 Sol | 82.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.7 Flash (thinking) qwen3-7-flash-epoch-qwen3-7-flash | Qwen3.7 Flash | 82.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-18 |
|---|
| GPT-5.6 Luna (low) gpt-5-6-luna-low | GPT-5.6 Luna | 82.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.7 Plus (no thinking) qwen-3-7-plus-epoch-qwen3-7-plus-none | Qwen3.7-Plus | 81.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| o3 (medium) o3-epoch-o3-2025-04-16-medium | o3 | 80.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Qwen3.7 Flash (no thinking) qwen3-7-flash-epoch-qwen3-7-flash-none | Qwen3.7 Flash | 80.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-18 |
|---|
| o4-mini (medium) o4-mini-epoch-o4-mini-2025-04-16-medium | o4-mini | 77.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.6 Terra (none) gpt-5-6-terra-epoch-gpt-5-6-terra-none | GPT-5.6 Terra | 77.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.5 (no thinking) gpt-5-5-epoch-gpt-5-5-none | GPT-5.5 | 77.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.4 nano (low) gpt-5-4-nano-epoch-gpt-5-4-nano-2026-03-17-low | GPT-5.4 nano | 72.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 71.7% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.1 (no thinking) gpt-5-1-epoch-gpt-5-1-2025-11-13-none | GPT-5.1 | 66.7% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.4 mini (none) gpt-5-4-mini-epoch-gpt-5-4-mini-2026-03-17-none | GPT-5.4 mini | 64.1% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.6 Luna (none) gpt-5-6-luna-epoch-gpt-5-6-luna-none | GPT-5.6 Luna | 63.6% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| GPT-5.4 nano (no thinking) gpt-5-4-nano-epoch-gpt-5-4-nano-2026-03-17-none | GPT-5.4 nano | 55.6% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-07 · checked 2026-08-17 |
|---|
| Gemini 3.1 Pro Preview (high) gemini-3-1-pro-preview-livebench-2026-06-25-high | Gemini 3.1 Pro Preview | 94.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 5 claude-opus-5-epoch-claude-opus-5 | Claude Opus 5 | 92.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| DeepSeek v4 (high) deepseek-v4-pro-high | DeepSeek V4 Pro | 90.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-18 |
|---|
| Gemini 3 Flash Preview (high) gemini-3-flash-epoch-gemini-3-flash-preview-high | Gemini 3 Flash | 89.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.5 Flash (low) gemini-3-5-flash-epoch-gemini-3-5-flash-low | Gemini 3.5 Flash | 88.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 4.8 (low) claude-opus-4-8-epoch-claude-opus-4-8-low | Claude Opus 4.8 | 88.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 4.6 (max) claude-opus-4-6-max | Claude Opus 4.6 | 88.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 5 (low) claude-opus-5-low | Claude Opus 5 | 87.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.6 Flash (low) gemini-3-6-flash-epoch-gemini-3-6-flash-low | Gemini 3.6 Flash | 86.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 4.7 (max) claude-opus-4-7-max | Claude Opus 4.7 | 86.4% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Fable 5 (max) claude-fable-5-max | Claude Fable 5 | 85.9% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Opus 4.8 (no thinking) claude-opus-4-8-epoch-claude-opus-4-8-none | Claude Opus 4.8 | 85.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.5 Flash-Lite (high) gemini-3-5-flash-lite-livebench-2026-06-25-high | Gemini 3.5 Flash-Lite | 83.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Fable 5 (high) claude-fable-5-high | Claude Fable 5 | 83.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.1 Flash-Lite (high) gemini-3-1-flash-lite-epoch-gemini-3-1-flash-lite-high | Gemini 3.1 Flash-Lite | 81.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Sonnet 5 (max) claude-sonnet-5-max | Claude Sonnet 5 | 80.3% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.6 (max) claude-sonnet-4-6-max | Claude Sonnet 4.6 | 78.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Claude Fable 5 (low) claude-fable-5-epoch-claude-fable-5-low | Claude Fable 5 | 78.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.5 Flash-Lite (low) gemini-3-5-flash-lite-epoch-gemini-3-5-flash-lite-low | Gemini 3.5 Flash-Lite | 75.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemma 4 31B IT (minimal) gemma-4-31b-epoch-gemma-4-31b-it-minimal | Gemma 4 31B | 75.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.5 Flash-Lite (minimal) gemini-3-5-flash-lite-epoch-gemini-3-5-flash-lite-minimal | Gemini 3.5 Flash-Lite | 74.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.1 Flash-Lite (low) gemini-3-1-flash-lite-epoch-gemini-3-1-flash-lite-low | Gemini 3.1 Flash-Lite | 74.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| Gemini 3.1 Flash-Lite (minimal) gemini-3-1-flash-lite-epoch-gemini-3-1-flash-lite-minimal | Gemini 3.1 Flash-Lite | 73.7% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-17 |
|---|
| DeepSeek v4 (no thinking) deepseek-v4-pro-epoch-deepseek-v4-pro-none | DeepSeek V4 Pro | 73.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-18 |
|---|
| gpt-oss-20b_high gpt-oss-20b-epoch-gpt-oss-20b-high | GPT-OSS 20B | 46.0% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-18 |
|---|
| gemma-3-27b-it gemma-3-27b-epoch-gemma-3-27b-it | Gemma 3 27B | 43.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-06 · checked 2026-08-18 |
|---|
| Inkling (xhigh) inkling-xhigh | Inkling | 88.3% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-18 |
|---|
| Qwen3.8 Max (xhigh) qwen-3-8-max-xhigh | Qwen3.8 Max | 92.7% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-04 · checked 2026-08-17 |
|---|
| Qwen3.8 Max (xhigh, default reasoning effort) qwen-3-8-max-xhigh | Qwen3.8 Max | 92.6% | CurrentQwen provider evaluation | Public reference · not scored provider-reported | Qwen3.8 Max official release ↗Observed 2026-08-03 · checked 2026-08-05 |
|---|
| Gemini 3.6 Flash (high) gemini-3-6-flash-high | Gemini 3.6 Flash | 94.1% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-02 · checked 2026-08-17 |
|---|
| DeepSeek V4 Flash 0731 (max) deepseek-v4-flash-0731-epoch-deepseek-v4-flash-0731-max | DeepSeek V4 Flash 0731 | 91.0% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-08-02 · checked 2026-08-17 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 94.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 94.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 93.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 93.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 93.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 92.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 92.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 92.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 92.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 92.7% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 92.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 92.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 92.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 91.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 91.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 91.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 90.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 90.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 90.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 90.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 90.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 89.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Interfaze Beta; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Interfaze Beta | 89.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 89.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 89.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling-Small; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling-Small | 89.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 89.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 89.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 89.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 88.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 88.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 88.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 88.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 88.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 87.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 87.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 87.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 87.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 87.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 87.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 87.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 87.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 87% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 86.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 86% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 85.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 85.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 85.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 85.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3-pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-pro | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 84.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 84.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 84.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 83.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 82.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 82.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 82.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 81.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 81.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 81.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 81% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 79.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 78.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 78.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 78.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 77.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 76.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 76.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 76.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 75.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 75.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 74.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 73.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 72.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 72.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 71.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant ZAYA1-8B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-8B | 71% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 68% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 67.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 66.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 66.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 63.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 63.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 62.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 61.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 58.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 57.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Thinking | 57.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 57.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant ZAYA1-74B-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-74B-Preview | 57.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 56.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 54.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 52.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 51.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 51.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 49.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 48.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 48.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 46.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Soofi S 30B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Soofi S 30B-A3B | 43.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 42.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 42.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 41.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Instruct | 40.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 40.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 37.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 29% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 27.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 26.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant LFM2.5-230M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-230M | 25.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 24.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 23.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 94.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 94.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 93.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 93.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 93.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 92.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 92.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 92.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 92.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 92.7% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 92.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 92.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 92.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 91.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 91.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 91.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 90.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 90.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 90.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 90.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 90.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 89.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Interfaze Beta; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Interfaze Beta | 89.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 89.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 89.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 89.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 89.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 89.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 88.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 88.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 88.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 88.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 88.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 87.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 87.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 87.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 87.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 87.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 87.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 87.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 87.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 87% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 86.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 86% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 85.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 85.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 85.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 85.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3-pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-pro | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 84.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 84.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 84.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 83.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 82.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 82.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 82.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 81.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 81.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 81.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 81% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 79.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 78.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 78.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 78.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 77.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 76.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 76.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 76.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 75.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 75.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 74.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 73.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 72.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 72.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 71.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant ZAYA1-8B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-8B | 71% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 68% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 67.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 66.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 66.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 63.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 63.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 62.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 61.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 58.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 57.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Thinking | 57.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 57.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant ZAYA1-74B-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-74B-Preview | 57.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 56.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 54.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 52.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 51.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 51.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 49.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 48.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 48.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 46.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Soofi S 30B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Soofi S 30B-A3B | 43.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 42.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 42.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 41.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Instruct | 40.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 40.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 37.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 29% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 27.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 26.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant LFM2.5-230M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-230M | 25.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 24.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 23.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Claude Opus 5 (max) claude-opus-5-max | Claude Opus 5 | 93.9% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-24 · checked 2026-08-17 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 95.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 94.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 94.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 93.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 93.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 93.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 92.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 92.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 92.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 92.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis GPQA Diamond independent evaluation. source label without registered configuration ID | Gemini 3.6 Flash | 92.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Gemini 3.6 Flash (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 92.7% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 92.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 92.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 92.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 91.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 91.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 91.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 90.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 90.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 90.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 90.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.3 | 90.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 Codex | 89.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Interfaze Beta; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Interfaze Beta | 89.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 89.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Hy3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 | 89.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-6-max | Claude Opus 4.6 | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.7 Code; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.7 Code | 89.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 89.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B (Reasoning); bulk export does not retain a complete upstream harness configuration. qwen3-5-397b-thinking | Qwen3.5 397B A17B | 89.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 89.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 89.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 88.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.7 | 88.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 88.5% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 88.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 88.1% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 87.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 87.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 | 87.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 87.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 87.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 87.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1 | 87.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 87.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Pro | 87% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 87% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5 Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 Thinking | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 86.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 86.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 86% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.1-Codex-Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.1-Codex-Max | 86% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 85.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 85.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 31B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 31B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 85.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (high); bulk export does not retain a complete upstream harness configuration. gpt-5-high | GPT-5 | 85.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast (Reasoning) | 85.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5-Turbo | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4 Fast (Reasoning); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4 Fast (Reasoning) | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3-pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-pro | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 84.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Pro | 84.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5 (medium); bulk export does not retain a complete upstream harness configuration. gpt-5-medium | GPT-5 | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 84.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 84.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 84.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 83.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis GPQA Diamond independent evaluation. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 83.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Gemini 3.5 Flash-Lite (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2-Omni; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2-Omni | 82.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3 | 82.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 82.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 81.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek-R1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek-R1 | 81.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3 Flash | 81.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 81% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4.1 Opus Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4.1 Opus Thinking | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5V-Turbo; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5V-Turbo | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 80.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 79.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 26B A4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 26B A4B | 79.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 12B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 12B Unified | 78.8% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant K-Exaone; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | K-Exaone | 78.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 120B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 120B | 78.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1 (Reasoning); bulk export does not retain a complete upstream harness configuration. deepseek-v3-1-reasoning-default | DeepSeek V3.1 | 77.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Small 4 (Reasoning); bulk export does not retain a complete upstream harness configuration. mistral-small-4-reasoning | Mistral Small 4 | 76.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2 | 76.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3 Max | 76.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Trinity-Large-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Thinking | 76.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 76.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano 30B | 75.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.2 | 75.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3.5 128B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3.5 128B | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o3-mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o3-mini | 74.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant o1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | o1 | 74.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sarvam 105B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 105B | 73.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3.1 | 73.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.5-Air; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.5-Air | 73.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 72.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron Ultra 253B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron Ultra 253B | 72.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok Code Fast 1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok Code Fast 1 | 72.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 72.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 71.2% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant ZAYA1-8B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-8B | 71% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-OSS 20B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-OSS 20B | 68.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 4 Sonnet; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 4 Sonnet | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 2.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 2.5 Flash | 68.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Large 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 3 | 68% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Maverick; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Maverick | 67.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 | 66.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 mini | 66.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.1 Fast; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.1 Fast | 63.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sarvam 30B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sarvam 30B | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Trinity-Large-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Trinity-Large-Preview | 63.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.6 | 63.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Exaone 4.0 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 32B | 62.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek R1 Distill Qwen 32B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek R1 Distill Qwen 32B | 61.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Ling 2.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ling 2.6 Flash | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 1.5 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.5 Pro | 58.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 4 Scout; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 4 Scout | 58.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Medium 3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Medium 3 | 57.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E4B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E4B | 57.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Thinking; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Thinking | 57.6% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Phi-4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Phi-4 | 57.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant ZAYA1-74B-Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | ZAYA1-74B-Preview | 57.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Solar Pro 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Solar Pro 2 | 56.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 55.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4o; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o | 54.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Llama 3.1 405B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Llama 3.1 405B | 51.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-8B-A1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-8B-A1B | 51.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4.1 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4.1 nano | 51.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nova Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nova Pro | 49.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 3 Opus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Opus | 48.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mistral Large 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mistral Large 2 | 48.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Soofi S 30B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Soofi S 30B-A3B | 43.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 3 27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 3 27B | 42.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-4o mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-4o mini | 42.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Exaone 4.0 1.2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Exaone 4.0 1.2B | 42.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen2.5 Coder 32B Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen2.5 Coder 32B Instruct | 41.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Mellum2-12B-A2.5B-Instruct; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Mellum2-12B-A2.5B-Instruct | 40.9% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemma 4 E2B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemma 4 E2B | 37.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude 3 Haiku; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude 3 Haiku | 37.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-VL-1.6B-Extract; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-VL-1.6B-Extract | 28.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-1B | 28.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 1.0 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 1.0 Pro | 27.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-1B | 26.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniCPM5-1B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniCPM5-1B | 26.3% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-H-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-H-350M | 25.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant LFM2.5-230M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | LFM2.5-230M | 25.4% | 2023Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Granite-4.0-350M; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Granite-4.0-350M | 20.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| o1 (low) o1-epoch-o1-2024-12-17-low | o1 | 74.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-20 · checked 2026-08-17 |
|---|
| GPT-5 (minimal) gpt-5-epoch-gpt-5-2025-08-07-minimal | GPT-5 | 71.7% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-20 · checked 2026-08-17 |
|---|
| o3-mini-2025-01-31_low o3-mini-epoch-o3-mini-2025-01-31-low | o3-mini | 68.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-20 · checked 2026-08-18 |
|---|
| GPQA Diamond; max reasoning; provider-published Kimi K3 evaluation. kimi-k3-max | Kimi K3 | 93.5% | CurrentProvider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Kimi K3 (Max) kimi-k3-max | Kimi K3 | 93.1% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-16 · checked 2026-08-17 |
|---|
| Muse Spark 1.1 (xhigh); Artificial Analysis independent evaluation configuration as published. muse-spark-1-1-xhigh | Muse Spark 1.1 | 89.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Muse Spark 1.1 (xhigh) analysis ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| deepseek-chat deepseek-v3-2-epoch-deepseek-chat | DeepSeek V3.2 | 71.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-16 · checked 2026-08-17 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Sol | 94.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Sol | 94.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-sol ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 94.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gemini-3-1-pro-preview ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Mythos 5 | 94.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.8 | 93.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 93.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.5 | 93.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.5 | 93.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | MiniMax M3 | 92.9% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for minimax-m3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Terra | 92.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 92.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Fable 5 | 92.6% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-fable-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Terra | 92.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-terra ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Max | 92.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Qwen3.7-Max | 92.3% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for qwen3-7-max ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Luna | 92.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 92.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.5 Flash | 92.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gemini-3-5-flash-medium ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Opus 4.8 | 92.0% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-opus-4-8 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.4 | 92.0% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.3-Codex | 91.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-3-codex ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.6 | 91.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5.2 | 91.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 91.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4.20 0309 v2 (Reasoning) source label without registered configuration ID | Grok 4.20 | 91.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Luna | 91.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-luna ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Sonnet 5 | 91.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-sonnet-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 90.4% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 90.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 90.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.3 | 90.1% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Grok 4.3 | 90.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 4.6 | 89.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 Codex (xhigh) gpt-5-2-codex-xhigh | GPT-5.2 Codex | 89.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (high) gpt-5-4-epoch-gpt-5-4-2026-03-05-high | GPT-5.4 | 89.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Kimi K2.7 Code source label without registered configuration ID | Kimi K2.7 Code | 89.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Flash | 89.4% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-flash ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 89.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (medium) gpt-5-4-epoch-gpt-5-4-2026-03-05-medium | GPT-5.4 | 88.9% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Pro | 88.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-pro ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 Max Preview source label without registered configuration ID | Qwen3.6 Max | 88.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 87.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 87.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 87.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 87% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nemotron 3 Ultra 550B A55B (Reasoning) source label without registered configuration ID | Nemotron 3 Ultra | 86.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 3.5 Flash (minimal) gemini-3-5-flash-epoch-gemini-3-5-flash-minimal | Gemini 3.5 Flash | 86.4% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Qwen3 Max Thinking source label without registered configuration ID | Qwen3 Max | 86.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5 | 86% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 85.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 27B (Reasoning) source label without registered configuration ID | Qwen3.5 27B | 85.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 122B A10B (Reasoning) source label without registered configuration ID | Qwen3.5 122B | 85.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 85.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5 source label without registered configuration ID | MiMo-V2.5 | 84.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.5 source label without registered configuration ID | MiniMax M2.5 | 84.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (low) gpt-5-4-low | GPT-5.4 | 84.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 84.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3-pro source label without registered configuration ID | o3-pro | 84.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 35B A3B (Reasoning) source label without registered configuration ID | Qwen3.5 35B | 84.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 84.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemma 4 31B | 84.3% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 27B (Reasoning) source label without registered configuration ID | Qwen3.6 27B | 84.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 84.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 Thinking source label without registered configuration ID | Kimi K2 Thinking | 83.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 83.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 83.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.1 source label without registered configuration ID | MiniMax M2.1 | 83.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 nano | 82.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 82.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 Omni Plus source label without registered configuration ID | Qwen3.5 Omni Plus | 82.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 80.9% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 80.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| NVIDIA Nemotron 3 Super 120B A12B (Reasoning) source label without registered configuration ID | Nemotron 3 Super | 80% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 (low) o3-epoch-o3-2025-04-16-low | o3 | 79.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Claude 4 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4 | 79.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 79.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 79.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 26B A4B (Reasoning) source label without registered configuration ID | Gemma 4 26B | 79.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 3 mini Reasoning (high) source label without registered configuration ID | Grok 3 mini | 79.1% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 78.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 78.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 78.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 77.7% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2 source label without registered configuration ID | MiniMax M2 | 77.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 3.7 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 3.7 | 77.2% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Small 4 | 76.9% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-small-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 0905 source label without registered configuration ID | Kimi K2 0905 | 76.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Command A+ source label without registered configuration ID | Command A+ | 76.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 12B (Reasoning) source label without registered configuration ID | Gemma 4 12B Unified | 75.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Ling-2.6-1T source label without registered configuration ID | Ling 2.6 1T | 75.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Medium 3.5 | 74.8% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o1 source label without registered configuration ID | o1 | 74.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 (none) gpt-5-4-epoch-gpt-5-4-2026-03-05-none | GPT-5.4 | 74.7% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-15 · checked 2026-08-17 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 72.9% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok Code Fast 1 source label without registered configuration ID | Grok Code Fast 1 | 72.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 71.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Large 3 | 68.0% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-large-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Haiku 4.5 | 64.7% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for claude-4-5-haiku ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| grok-4.20-0309-reasoning grok-4-20-epoch-grok-4-20-0309-reasoning | Grok 4.20 | 89.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.6 (high) claude-sonnet-4-6-non-reasoning-high | Claude Sonnet 4.6 | 83.3% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.6 (medium) claude-sonnet-4-6-livebench-2026-06-25-medium | Claude Sonnet 4.6 | 83.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-17 |
|---|
| GPT-5.2 (none) gpt-5-2-epoch-gpt-5-2-2025-12-11-none | GPT-5.2 | 73.2% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-17 |
|---|
| GPT-5 nano (low) gpt-5-nano-epoch-gpt-5-nano-2025-08-07-low | GPT-5 nano | 57.6% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-18 |
|---|
| GPT-5 nano (low) gpt-5-nano-epoch-gpt-5-nano-2025-08-07-minimal | GPT-5 nano | 48.5% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-13 · checked 2026-08-18 |
|---|
| o4-mini (low) o4-mini-epoch-o4-mini-2025-04-16-low | o4-mini | 75.3% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-11 · checked 2026-08-17 |
|---|
| GPT-5.6 Sol (max) gpt-5-6-sol-max | GPT-5.6 Sol | 93.5% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-09 · checked 2026-08-17 |
|---|
| GPT-5.6 Terra (max) gpt-5-6-terra-max | GPT-5.6 Terra | 93.3% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-09 · checked 2026-08-17 |
|---|
| GPT-5.6 Luna (max) gpt-5-6-luna-max | GPT-5.6 Luna | 91.6% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-09 · checked 2026-08-17 |
|---|
| Grok 4.5 (high) grok-4-5-aa-2-high | Grok 4.5 | 93.4% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-08 · checked 2026-08-17 |
|---|
| tencent/Hy3 instruct weights; configuration as recorded on the Hugging Face evaluation results entry. source label without registered configuration ID | Hy3 | 90.4% | CurrentProvider / community evaluation result attached on Hugging Face | Score input provider-reported | Tencent Hy3 model card on Hugging Face ↗Observed 2026-07-06 · checked 2026-07-16 |
|---|
| Claude Sonnet 5 (xhigh) claude-sonnet-5-xhigh | Claude Sonnet 5 | 90.5% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-07-01 · checked 2026-08-17 |
|---|
| LongCat-2.0 (in-house harness) longcat-2-0-default | LongCat-2.0 | 88.9% | DiamondLongCat unified in-house harness | Score input provider-reported | LongCat-2.0 model card ↗Observed 2026-06-30 · checked 2026-08-15 |
|---|
| GLM-5.2 (max) glm-5-2-max | GLM-5.2 | 91.9% | 1.0.11Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-06-24 · checked 2026-08-17 |
|---|
| grok-4.3_high grok-4-3-epoch-grok-4-3-high | Grok 4.3 | 88.8% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-06-17 · checked 2026-08-17 |
|---|
| DeepSeek v4 (max) deepseek-v4-pro-max | DeepSeek V4 Pro | 89.6% | 1.0.11Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-06-16 · checked 2026-08-18 |
|---|
| Claude Opus 4.8 (max) claude-opus-4-8-max | Claude Opus 4.8 | 91.0% | 1.0.10Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-06-07 · checked 2026-08-17 |
|---|
| Gemini 3.5 Flash (high) gemini-3-5-flash-high | Gemini 3.5 Flash | 92.8% | 1.0.9Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-05-22 · checked 2026-08-17 |
|---|
| GPT-5.5 (low) gpt-5-5-low | GPT-5.5 | 90.7% | 1.0.8Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-05-05 · checked 2026-08-17 |
|---|
| Kimi K2.6 kimi-k2-6-epoch-kimi-k2-6 | Kimi K2.6 | 90.8% | 1.0.8Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-05-01 · checked 2026-08-17 |
|---|
| GPT-5.5 (xhigh) gpt-5-5-xhigh | GPT-5.5 | 94.0% | 1.0.8Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-04-24 · checked 2026-08-17 |
|---|
| GPT-5.5 Pro (xhigh) gpt-5-5-pro | GPT-5.5 | 93.9% | 1.0.8Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-04-24 · checked 2026-08-17 |
|---|
| Claude Opus 4.7 (xhigh) claude-opus-4-7-livebench-2026-06-25-xhigh | Claude Opus 4.7 | 90.2% | 1.0.7Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-04-17 · checked 2026-08-17 |
|---|
| GPT-5.4 mini (high) gpt-5-4-mini-epoch-gpt-5-4-mini-2026-03-17-high | GPT-5.4 mini | 83.6% | 1.0.6Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-04-15 · checked 2026-08-17 |
|---|
| GPT-5.4 nano (high) gpt-5-4-nano-epoch-gpt-5-4-nano-2026-03-17-high | GPT-5.4 nano | 78.5% | 1.0.6Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-04-14 · checked 2026-08-17 |
|---|
| GPT-5.4 Pro (xhigh) gpt-5-4-pro-xhigh | GPT-5.4 | 94.6% | 1.0.6Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-03-20 · checked 2026-08-18 |
|---|
| GPT-5.4 (xhigh) gpt-5-4-xhigh | GPT-5.4 | 93.3% | 1.0.6Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2026-03-06 · checked 2026-08-17 |
|---|
| Gemini 3.1 Pro Preview gemini-3-1-pro-preview-epoch-gemini-3-1-pro-preview | Gemini 3.1 Pro Preview | 94.1% | 1.0.5Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-20 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.6 (32k thinking) claude-sonnet-4-6-epoch-claude-sonnet-4-6-32k | Claude Sonnet 4.6 | 87.4% | 1.0.6Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-20 · checked 2026-08-17 |
|---|
| GLM-5 glm-5-epoch-glm-5 | GLM-5 | 87.8% | 1.0.5Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-12 · checked 2026-08-17 |
|---|
| Claude Opus 4.6 (32k thinking) claude-opus-4-6-epoch-claude-opus-4-6-32k | Claude Opus 4.6 | 90.5% | 1.0.4Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-06 · checked 2026-08-17 |
|---|
| Claude Opus 4.6 (64k thinking) claude-opus-4-6-epoch-claude-opus-4-6-64k | Claude Opus 4.6 | 88.8% | 1.0.4Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-06 · checked 2026-08-17 |
|---|
| Kimi K2.5 (Fireworks) kimi-k2-5-epoch-fireworks-kimi-k2p5 | Kimi K2.5 | 87.6% | 1.0.4Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-02-02 · checked 2026-08-17 |
|---|
| GLM-4.7 glm-4-7-epoch-glm-4-7 | GLM-4.7 | 83.3% | 1.0.4Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2026-01-29 · checked 2026-08-17 |
|---|
| Kimi K2.5 thinking configuration kimi-k2-5-thinking | Kimi K2.5 | 87.6% | diamondKimi K2.5 official model-card evaluation | Source carrier · no additional scoring weight source-checked | Kimi K2.5 official model card and evaluation table ↗Observed 2026-01-27 · checked 2026-08-29 |
|---|
| Gemini 3 Flash Preview gemini-3-flash-epoch-gemini-3-flash-preview | Gemini 3 Flash | 83.2% | 1.0.3Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-17 · checked 2026-08-17 |
|---|
| DeepSeek-V3.2 (Thinking) deepseek-v3-2-epoch-deepseek-reasoner | DeepSeek V3.2 | 83.4% | 1.0.3Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-16 · checked 2026-08-17 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 91.4% | 1.0.3Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-13 · checked 2026-08-17 |
|---|
| GPT-5.2 (high) gpt-5-2-livebench-2026-06-25-high | GPT-5.2 | 88.2% | 1.0.3Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-11 · checked 2026-08-17 |
|---|
| GPT-5.2 (medium) gpt-5-2-medium | GPT-5.2 | 87.9% | 1.0.3Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-11 · checked 2026-08-17 |
|---|
| GPT-5.2 (low) gpt-5-2-epoch-gpt-5-2-2025-12-11-low | GPT-5.2 | 82.7% | 1.0.3Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-11 · checked 2026-08-17 |
|---|
| gpt-oss-120b (high) gpt-oss-120b-high | GPT-OSS 120B | 75.8% | 1.0.3Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2025-12-11 · checked 2026-08-18 |
|---|
| Claude Opus 4.5 (16k thinking) claude-opus-4-5-epoch-claude-opus-4-5-20251101-16k | Claude Opus 4.5 | 85.5% | 1.0.2Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-25 · checked 2026-08-17 |
|---|
| Claude Opus 4.5 (32k thinking) claude-opus-4-5-epoch-claude-opus-4-5-20251101-32k | Claude Opus 4.5 | 86.0% | 1.0.2Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-24 · checked 2026-08-17 |
|---|
| Claude Opus 4.5 (no thinking) claude-opus-4-5-epoch-claude-opus-4-5-20251101 | Claude Opus 4.5 | 80.7% | 1.0.2Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-24 · checked 2026-08-17 |
|---|
| Gemini 3 Pro Preview gemini-3-pro-epoch-gemini-3-pro-preview | Gemini 3 Pro Preview | 92.6% | 1.0.2Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-19 · checked 2026-08-18 |
|---|
| GPT-5.1 (medium) gpt-5-1-epoch-gpt-5-1-2025-11-13-medium | GPT-5.1 | 85.0% | 1.0.1Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-17 · checked 2026-08-17 |
|---|
| Gemini 2.5 Pro gemini-2-5-pro-epoch-gemini-2-5-pro | Gemini 2.5 Pro | 85.3% | 1.0.1Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-16 · checked 2026-08-17 |
|---|
| GPT-5.1 (high) gpt-5-1-high | GPT-5.1 | 87.6% | 1.0.1Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-13 · checked 2026-08-17 |
|---|
| Kimi K2 Thinking Turbo kimi-k2-thinking-epoch-kimi-k2-thinking-turbo | Kimi K2 Thinking | 84.2% | 1.0.1Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-11-11 · checked 2026-08-17 |
|---|
| GPT-5 mini (high) gpt-5-mini-high | GPT-5 mini | 75% | 1.0.1Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-30 · checked 2026-08-17 |
|---|
| GPT-5 nano (high) gpt-5-nano-high | GPT-5 nano | 69.4% | 1.0.1Epoch AI benchmark runner | Score input source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-30 · checked 2026-08-18 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 86.2% | 1.0.1Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-29 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.5 (59k thinking) claude-sonnet-4-5-epoch-claude-sonnet-4-5-20250929-59k | Claude Sonnet 4.5 | 82.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-28 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.5 (16k thinking) claude-sonnet-4-5-epoch-claude-sonnet-4-5-20250929-16k | Claude Sonnet 4.5 | 78.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-28 · checked 2026-08-17 |
|---|
| Claude Haiku 4.5 (32k thinking) claude-haiku-4-5-epoch-claude-haiku-4-5-20251001-32k | Claude Haiku 4.5 | 71.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-22 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.5 (32k thinking) claude-sonnet-4-5-epoch-claude-sonnet-4-5-20250929-32k | Claude Sonnet 4.5 | 81.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-21 · checked 2026-08-17 |
|---|
| claude-haiku-4-5-20251001 claude-haiku-4-5-epoch-claude-haiku-4-5-20251001 | Claude Haiku 4.5 | 60.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-16 · checked 2026-08-17 |
|---|
| Qwen3 Max qwen3-max-epoch-qwen3-max-2025-09-23 | Qwen3 Max | 72.6% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-10-06 · checked 2026-08-17 |
|---|
| Claude Sonnet 4.5 (no thinking) claude-sonnet-4-5-epoch-claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 | 73.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-09-29 · checked 2026-08-17 |
|---|
| GPT-5 (medium) gpt-5-medium | GPT-5 | 85.4% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-07 · checked 2026-08-17 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 71.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-07 · checked 2026-08-17 |
|---|
| GPT-5 nano (medium) gpt-5-nano-medium | GPT-5 nano | 67.4% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-07 · checked 2026-08-18 |
|---|
| Claude Opus 4.1 (16k thinking) claude-opus-4-1-epoch-claude-opus-4-1-20250805-16k | Claude Opus 4.1 | 77.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-05 · checked 2026-08-18 |
|---|
| Claude Opus 4.1 (27k thinking) claude-opus-4-1-epoch-claude-opus-4-1-20250805-27k | Claude Opus 4.1 | 76.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-05 · checked 2026-08-18 |
|---|
| Claude Opus 4.1 claude-opus-4-1-epoch-claude-opus-4-1-20250805 | Claude Opus 4.1 | 73.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-08-05 · checked 2026-08-18 |
|---|
| Gemini 2.5 Pro Preview (Jun 2025) gemini-2-5-pro-epoch-gemini-2-5-pro-preview-06-05 | Gemini 2.5 Pro | 84.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-06-05 · checked 2026-08-17 |
|---|
| Claude Sonnet 4 (59k thinking) claude-sonnet-4-epoch-claude-sonnet-4-20250514-59k | Claude Sonnet 4 | 79.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-26 · checked 2026-08-18 |
|---|
| Grok 3 (beta) grok-3-epoch-grok-3-beta | Grok 3 | 75.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-26 · checked 2026-08-18 |
|---|
| grok-3-mini-beta_high grok-3-mini-epoch-grok-3-mini-beta-high | Grok 3 mini | 75.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-26 · checked 2026-08-17 |
|---|
| DeepSeek-R1 deepseek-r1-epoch-deepseek-r1 | DeepSeek-R1 | 71.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-26 · checked 2026-08-18 |
|---|
| Claude Sonnet 4 (32k thinking) claude-sonnet-4-epoch-claude-sonnet-4-20250514-32k | Claude Sonnet 4 | 78.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-22 · checked 2026-08-18 |
|---|
| Claude Opus 4 (16k thinking) claude-opus-4-epoch-claude-opus-4-20250514-16k | Claude Opus 4 | 76.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-22 · checked 2026-08-18 |
|---|
| Claude Sonnet 4 (16k thinking) claude-sonnet-4-epoch-claude-sonnet-4-20250514-16k | Claude Sonnet 4 | 75.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-22 · checked 2026-08-18 |
|---|
| Claude Opus 4 claude-opus-4-epoch-claude-opus-4-20250514 | Claude Opus 4 | 69.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-22 · checked 2026-08-18 |
|---|
| Claude Sonnet 4 claude-sonnet-4-epoch-claude-sonnet-4-20250514 | Claude Sonnet 4 | 66.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-22 · checked 2026-08-18 |
|---|
| mistral-medium-2505 mistral-medium-3-epoch-mistral-medium-2505 | Mistral Medium 3 | 59.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-05-07 · checked 2026-08-18 |
|---|
| o3 (high) o3-epoch-o3-2025-04-16-high | o3 | 81.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-16 · checked 2026-08-17 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 79.6% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-16 · checked 2026-08-17 |
|---|
| GPT-4.1 gpt-4-1-epoch-gpt-4-1-2025-04-14 | GPT-4.1 | 66.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-14 · checked 2026-08-18 |
|---|
| GPT-4.1 mini gpt-4-1-mini-epoch-gpt-4-1-mini-2025-04-14 | GPT-4.1 mini | 65.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-14 · checked 2026-08-18 |
|---|
| GPT-4.1 nano gpt-4-1-nano-epoch-gpt-4-1-nano-2025-04-14 | GPT-4.1 nano | 48.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-14 · checked 2026-08-18 |
|---|
| grok-3-mini-beta_low grok-3-mini-epoch-grok-3-mini-beta-low | Grok 3 mini | 76.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-10 · checked 2026-08-17 |
|---|
| Llama 4 Maverick (FP8) llama-4-maverick-epoch-llama-4-maverick-17b-128e-instruct-fp8 | Llama 4 Maverick | 67.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-08 · checked 2026-08-18 |
|---|
| Llama-4-Scout-17B-16E-Instruct llama-4-scout-epoch-llama-4-scout-17b-16e-instruct | Llama 4 Scout | 51.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-08 · checked 2026-08-18 |
|---|
| Qwen Turbo qwen-turbo-epoch-qwen-turbo-2024-11-01 | Qwen2.5 Turbo | 41.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-07 · checked 2026-08-18 |
|---|
| Qwen2.5-Max qwen-2-5-max-epoch-qwen-max-2025-01-25 | Qwen2.5 Max | 56.1% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-04-01 · checked 2026-08-18 |
|---|
| mistral-small-2503 mistral-small-3-1-epoch-mistral-small-2503 | Mistral Small 3.1 | 47.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-03-18 · checked 2026-08-18 |
|---|
| Claude 3.5 Haiku (Oct 2024) claude-3-5-haiku-epoch-claude-3-5-haiku-20241022 | Claude 3.5 Haiku | 38.1% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-03-12 · checked 2026-08-18 |
|---|
| gemini-1.5-flash-8b-001 gemini-1-5-flash-8b-epoch-gemini-1-5-flash-8b-001 | Gemini 1.5 Flash-8B | 33.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-03-12 · checked 2026-08-18 |
|---|
| DeepSeek-R1-Distill-Llama-70B deepseek-r1-distill-llama-70b-epoch-deepseek-r1-distill-llama-70b | DeepSeek R1 Distill Llama 70B | 55.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-03-10 · checked 2026-08-18 |
|---|
| DeepSeek-R1-Distill-Qwen-14B deepseek-r1-distill-qwen-14b-epoch-deepseek-r1-distill-qwen-14b | DeepSeek R1 Distill Qwen 14B | 44.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-03-10 · checked 2026-08-18 |
|---|
| GPT-4.5 Preview (Feb 2025) gpt-4-5-epoch-gpt-4-5-preview-2025-02-27 | GPT-4.5 (Preview) | 68.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-28 · checked 2026-08-18 |
|---|
| mistral-large-2411 mistral-large-2-epoch-mistral-large-2411 | Mistral Large 2 | 51.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-25 · checked 2026-08-18 |
|---|
| o3-mini (high) o3-mini-high | o3-mini | 77.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-13 · checked 2026-08-18 |
|---|
| o1 (high) o1-epoch-o1-2024-12-17-high | o1 | 76.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-13 · checked 2026-08-17 |
|---|
| o1-mini (high) o1-mini-epoch-o1-mini-2024-09-12-high | o1-mini | 62.4% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-13 · checked 2026-08-18 |
|---|
| Gemini 2.0 Flash (Feb 2025) gemini-2-0-flash-epoch-gemini-2-0-flash-001 | Gemini 2.0 Flash (Feb '25) | 64.1% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-06 · checked 2026-08-18 |
|---|
| GPT-4o (Nov 2024) gpt-4o-epoch-gpt-4o-2024-11-20 | GPT-4o | 47.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-02-05 · checked 2026-08-18 |
|---|
| o3-mini (medium) o3-mini-epoch-o3-mini-2025-01-31-medium | o3-mini | 74.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-31 · checked 2026-08-18 |
|---|
| phi-4 phi-4-epoch-phi-4 | Phi-4 | 56.1% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-31 · checked 2026-08-18 |
|---|
| mistral-small-2501 mistral-small-3-epoch-mistral-small-2501 | Mistral Small 3 | 45.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-30 · checked 2026-08-18 |
|---|
| o1 (medium) o1-epoch-o1-2024-12-17-medium | o1 | 75.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-17 |
|---|
| o1-mini (medium) o1-mini-epoch-o1-mini-2024-09-12-medium | o1-mini | 59.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| gemini-1.5-pro-002 gemini-1-5-pro-epoch-gemini-1-5-pro-002 | Gemini 1.5 Pro | 57.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| DeepSeek-V3 deepseek-v3-epoch-deepseek-v3 | DeepSeek V3 | 56.5% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Claude 3.5 Sonnet (Jun 2024) claude-3-5-sonnet-epoch-claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet | 54.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Llama-3.1-405B-Instruct llama-3-1-405b-epoch-llama-3-1-405b-instruct | Llama 3.1 405B | 50.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| o1-preview o1-preview-epoch-o1-preview-2024-09-12 | o1-preview | 50.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| GPT-4o (Aug 2024) gpt-4o-epoch-gpt-4o-2024-08-06 | GPT-4o | 49.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Qwen2.5-72B qwen2-5-72b-epoch-qwen2-5-72b-instruct | Qwen2.5-72B | 49.1% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| mistral-large-2407 mistral-large-2-epoch-mistral-large-2407 | Mistral Large 2 | 49.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| GPT-4o (May 2024) gpt-4o-epoch-gpt-4o-2024-05-13 | GPT-4o | 48.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| gemini-1.5-flash-002 gemini-1-5-flash-epoch-gemini-1-5-flash-002 | Gemini 1.5 Flash (Sep '24) | 47.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Claude 3 Opus claude-3-opus-epoch-claude-3-opus-20240229 | Claude 3 Opus | 47.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| gemini-1.5-pro-001 gemini-1-5-pro-epoch-gemini-1-5-pro-001 | Gemini 1.5 Pro | 45.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Claude 3 Sonnet claude-3-sonnet-epoch-claude-3-sonnet-20240229 | Claude 3 Sonnet | 40.6% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Meta-Llama-3-70B-Instruct llama-3-70b-epoch-meta-llama-3-70b-instruct | Llama 3 70B | 40.6% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Gemini 1.5 Flash (May 2024) gemini-1-5-flash-epoch-gemini-1-5-flash-001 | Gemini 1.5 Flash (Sep '24) | 40.4% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| mistral-large-2402 mistral-large-epoch-mistral-large-2402 | Mistral Large (Feb '24) | 38.8% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| gpt-4o-mini-2024-07-18 gpt-4o-mini-epoch-gpt-4o-mini-2024-07-18 | GPT-4o mini | 37.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| Claude 3 Haiku claude-3-haiku-epoch-claude-3-haiku-20240307 | Claude 3 Haiku | 36.3% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| claude-2.0 claude-2-epoch-claude-2-0 | Claude 2.0 | 34.7% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| gemini-1.0-pro-001 gemini-1-0-pro-epoch-gemini-1-0-pro-001 | Gemini 1.0 Pro | 34.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| claude-2.1 claude-21-epoch-claude-2-1 | Claude 2.1 | 33.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| DBRX (instruct) dbrx-instruct-epoch-dbrx-instruct | DBRX Instruct | 32.9% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| GPT-3.5 Turbo (Nov 2023) gpt-35-turbo-epoch-gpt-3-5-turbo-1106 | GPT-3.5 Turbo | 28.0% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|
| GPT-3.5 Turbo (Jan 2024) gpt-35-turbo-epoch-gpt-3-5-turbo-0125 | GPT-3.5 Turbo | 27.2% | 1.0.0Epoch AI benchmark runner | Public reference · not scored source-checked | Epoch AI permanent refresh source ↗Observed 2025-01-27 · checked 2026-08-18 |
|---|