| Qwen3.8-Flash-Next (Artificial Analysis independent run) qwen-3-8-flash-next-aa-unspecified | Qwen3.8-Flash-Next | 86.1% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Public reference · not scored independently-verified | Qwen3.8-Flash-Next individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) (Artificial Analysis independent run) claude-opus-4-7-max | Claude Opus 4.7 | 83.1% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GLM-5.2 (max) (Artificial Analysis independent run) glm-5-2-max | GLM-5.2 | 77.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 397B A17B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-397b-thinking | Qwen3.5 397B A17B | 51.3% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Qwen3.5 397B A17B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Medium 3.5 (Artificial Analysis completed independent run) mistral-medium-3-5-128b-aa-reasoning-default | Mistral Medium 3.5 128B | 50.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Mistral Medium 3.5 current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| LongCat 2.0 (Artificial Analysis independent run) longcat-2-0-default | LongCat-2.0 | 50.2% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | LongCat 2.0 individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.5 122B A10B (Reasoning) (Artificial Analysis completed independent run) qwen3-5-122b-a10b-aa-reasoning-default | Qwen3.5-122B-A10B | 47.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Qwen3.5 122B A10B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Kimi K2.5 (Reasoning) (Artificial Analysis completed independent run) kimi-k2-5-thinking | Kimi K2.5 | 45.7% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Kimi K2.5 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Qwen3.6 35B A3B (Reasoning) (Artificial Analysis completed independent run) qwen3-6-35b-a3b-aa-reasoning-default | Qwen3.6-35B-A3B | 44.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Qwen3.6 35B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 4 26B A4B (Reasoning) (Artificial Analysis completed independent run) gemma-4-26b-a4b-aa-reasoning-default | Gemma 4 26B A4B | 39.0% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Gemma 4 26B A4B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Mistral Small 4 (Reasoning) (Artificial Analysis completed independent run) mistral-small-4-reasoning | Mistral Small 4 | 21.0% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Mistral Small 4 (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Trinity Large Thinking (Artificial Analysis completed independent run) trinity-large-thinking-aa-reasoning-default | Trinity-Large-Thinking | 20.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Trinity Large Thinking current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 mini (Artificial Analysis completed independent run) gpt-4-1-mini-epoch-gpt-4-1-mini-2025-04-14 | GPT-4.1 mini | 10.1% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-4.1 mini current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Maverick (Artificial Analysis completed independent run) llama-4-maverick-aa-non-reasoning-default | Llama 4 Maverick | 7.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Llama 4 Maverick current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) (Artificial Analysis completed independent run) nemotron-3-nano-30b-aa-reasoning-default | Nemotron 3 Nano 30B | 6.7% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Gemma 3 27B Instruct (Artificial Analysis completed independent run) gemma-3-27b-aa-non-reasoning-default | Gemma 3 27B | 4.5% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Gemma 3 27B Instruct current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| GPT-4.1 nano (Artificial Analysis completed independent run) gpt-4-1-nano-epoch-gpt-4-1-nano-2025-04-14 | GPT-4.1 nano | 3.7% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-4.1 nano current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Llama 4 Scout (Artificial Analysis completed independent run) llama-4-scout-aa-non-reasoning-default | Llama 4 Scout | 3.7% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Llama 4 Scout current model and individual-evaluation page ↗Observed 2026-08-29 · checked 2026-08-29 |
|---|
| Claude Opus 5 (max; Artificial Analysis independent run) claude-opus-5-max | Claude Opus 5 | 89.1% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Claude Opus 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Grok 4.6 (high; Artificial Analysis independent run) grok-4-6-high | Grok 4.6 | 88.4% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Grok 4.6 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Sol (max; Artificial Analysis independent run) gpt-5-6-sol-max | GPT-5.6 Sol | 88.0% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-5.6 Sol individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Terra (max; Artificial Analysis independent run) gpt-5-6-terra-max | GPT-5.6 Terra | 88.0% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-5.6 Terra individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Kimi K3 (max; Artificial Analysis independent run) kimi-k3-max | Kimi K3 | 85.0% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Kimi K3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Fable 5 (max; Artificial Analysis independent run) claude-fable-5-max | Claude Fable 5 | 84.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Claude Fable 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Opus 4.8 (max; Artificial Analysis independent run) claude-opus-4-8-max | Claude Opus 4.8 | 84.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Claude Opus 4.8 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3-Flash (max; Artificial Analysis independent run) glm-5-3-flash-max | GLM-5.3-Flash | 84.3% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GLM-5.3-Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.5 (xhigh; Artificial Analysis independent run) gpt-5-5-xhigh | GPT-5.5 | 84.3% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-5.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GLM-5.3 (max; Artificial Analysis independent run) glm-5-3-max | GLM-5.3 | 83.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GLM-5.3 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Grok 4.5 (high; Artificial Analysis independent run) grok-4-5-aa-2-high | Grok 4.5 | 81.6% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Grok 4.5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.6 Luna (max; Artificial Analysis independent run) gpt-5-6-luna-max | GPT-5.6 Luna | 80.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-5.6 Luna individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Claude Sonnet 5 (max; Artificial Analysis independent run) claude-sonnet-5-max | Claude Sonnet 5 | 80.5% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Claude Sonnet 5 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.2 (xhigh; Artificial Analysis independent run) muse-spark-1-2-xhigh | Muse Spark 1.2 | 80.1% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Muse Spark 1.2 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Pro 0813 (max; Artificial Analysis independent run) deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 78.7% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Public reference · not scored independently-verified | DeepSeek V4 Pro 0813 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.7 Flash (medium; Artificial Analysis independent run) gemini-3-7-flash-medium | Gemini 3.7 Flash | 78.3% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Gemini 3.7 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| GPT-5.4 (xhigh; Artificial Analysis independent run) gpt-5-4-xhigh | GPT-5.4 | 78.3% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | GPT-5.4 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Muse Spark 1.1 (xhigh; Artificial Analysis independent run) muse-spark-1-1-xhigh | Muse Spark 1.1 | 77.9% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Muse Spark 1.1 individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| Gemini 3.6 Flash (high; Artificial Analysis independent run) gemini-3-6-flash-high | Gemini 3.6 Flash | 77.5% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Score input independently-verified | Gemini 3.6 Flash individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek V4 Flash Vision Exp (max; Artificial Analysis independent run) deepseek-v4-flash-vision-exp-max-harness | DeepSeek V4 Flash Vision Exp | 74.2% | 2.1Artificial Analysis Terminal-Bench v2.1 in e2b | Public reference · not scored independently-verified | DeepSeek V4 Flash Vision Exp individual evaluations ↗Observed 2026-08-27 · checked 2026-08-27 |
|---|
| DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0 deepseek-v4-flash-vision-exp-max-harness | DeepSeek V4 Flash Vision Exp | 83.9% | 2.1DeepSeek Harness Minimal Mode | Public reference · not scored provider-reported | DeepSeek-V4-Flash-Vision-Exp release and provider evaluation ↗Observed 2026-08-21 · checked 2026-08-24 |
|---|
| Claude Opus 4.6 Max as published by Qwen claude-opus-4-6-max | Claude Opus 4.6 | 78.2% | 2.1Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 73% | 2.1Qwen3.8-27B official text table | Score input provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 64% | 2.1Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 63.4% | 2.1Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| GPT-5.6 Sol as published by Z.AI gpt-5-6-sol-max | GPT-5.6 Sol | 88.8% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Kimi K3 as published by Z.AI kimi-k3-max | Kimi K3 | 88.3% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 88.2% | 2.1Z.AI GLM-5.3 official comparison table | Score input provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Fable 5 w/ fallback as published by Z.AI claude-fable-5-max | Claude Fable 5 | 88% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| DeepSeek V4 Pro 0813 as published by Z.AI deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 87.9% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Qwen3.8 Max as published by Z.AI qwen-3-8-max-xhigh | Qwen3.8 Max | 86.6% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Opus 4.8 as published by Z.AI claude-opus-4-8-max | Claude Opus 4.8 | 85% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.2 as published by Z.AI glm-5-2-max | GLM-5.2 | 81% | 2.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Kimi K3 kimi-k3-deepseek-0813-release-unspecified | Kimi K3 | 88.3% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Fable 5 (with fallback) claude-fable-5-deepseek-0813-release-unspecified | Claude Fable 5 | 88% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Pro 0813 deepseek-v4-pro-0813-deepseek-0813-release-unspecified | DeepSeek V4 Pro 0813 | 87.9% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GPT-5.6 Terra gpt-5-6-terra-max | GPT-5.6 Terra | 87.4% | 2.1Terminal-Bench 2.1 default Terminus 2 agent harness | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.7 Flash gemini-3-7-flash-medium | Gemini 3.7 Flash | 85.8% | 2.1Terminal-Bench 2.1 default Terminus 2 agent harness | Score input provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Opus 4.8 claude-opus-4-8-deepseek-0813-release-unspecified | Claude Opus 4.8 | 85% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Muse Spark 1.2 muse-spark-1-2-xhigh | Muse Spark 1.2 | 82.9% | 2.1Terminal-Bench 2.1 default Terminus 2 agent harness | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Flash 0731 deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 82.7% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Source carrier · no additional scoring weight provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GLM-5.2 glm-5-2-deepseek-0813-release-unspecified | GLM-5.2 | 81% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Sonnet 5 claude-sonnet-5-max | Claude Sonnet 5 | 80.4% | 2.1Terminal-Bench 2.1 default Terminus 2 agent harness | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.6 Flash gemini-3-6-flash-high | Gemini 3.6 Flash | 78% | 2.1Terminal-Bench 2.1 default Terminus 2 agent harness | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Pro Preview deepseek-v4-pro-deepseek-0813-release-unspecified | DeepSeek V4 Pro | 72.1% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Flash Preview deepseek-v4-flash-deepseek-0813-release-unspecified | DeepSeek V4 Flash | 61.8% | 2.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16; NVIDIA release evaluation; temperature=1.0; top_p=0.95 nemotron-3-5-lightning-30b-a3b-default | NVIDIA Nemotron 3.5 Lightning 30B-A3B | 24.6% | 2.1NVIDIA NeMo Gym / NeMo Evaluator SDK consistent release harness | Score input provider-reported | NVIDIA Nemotron 3.5 Lightning 30B-A3B BF16 model card ↗Observed 2026-08-11 · checked 2026-08-12 |
|---|
| Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64 muse-glimmer-30b-high | Muse Glimmer 30B | 51.7% | 2.1Artificial Analysis Terminal-Bench 2.1 evaluation in an E2B sandbox | Public reference · not scored provider-reported | Muse Glimmer Evaluation Methodology ↗Observed 2026-08-10 · checked 2026-08-10 |
|---|
| Claude Opus 5 (reasoning configuration not stated in chart) claude-opus-5-meta-muse-12-release-unspecified | Claude Opus 5 | 86.7% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.2 (xhigh) muse-spark-1-2-xhigh | Muse Spark 1.2 | 82.9% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Public reference · not scored provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.2 (xhigh) with Muse Code muse-spark-1-2-xhigh | Muse Spark 1.2 | 82.9% | 2.1Muse Code evaluation on all 89 Terminal-Bench 2.1 tasks in isolated Daytona sandboxes | Public reference · not scored provider-reported | Muse Spark 1.2 Evaluation Methodology ↗Observed 2026-08-05 · checked 2026-08-05 |
|---|
| GPT-5.6 Terra (reasoning configuration not stated in chart) gpt-5-6-terra-meta-muse-12-release-unspecified | GPT-5.6 Terra | 81.8% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Grok 4.5 (reasoning configuration not stated in chart) grok-4-5-meta-muse-12-release-unspecified | Grok 4.5 | 81.6% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Gemini 3.6 Flash (reasoning configuration not stated in chart) gemini-3-6-flash-meta-muse-12-release-unspecified | Gemini 3.6 Flash | 78.9% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.1 (reasoning configuration not stated in chart) muse-spark-1-1-meta-muse-12-release-unspecified | Muse Spark 1.1 | 76.2% | 2.1Meta release chart for Terminal-Bench 2.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Qwen3.8 Max (xhigh, default reasoning effort) qwen-3-8-max-xhigh | Qwen3.8 Max | 86.6% | 2.1Claude Code on Terminal-Bench 2.1 | Public reference · not scored provider-reported | Qwen3.8 Max official release ↗Observed 2026-08-03 · checked 2026-08-05 |
|---|
| DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0 deepseek-v4-flash-0731-max-harness | DeepSeek V4 Flash 0731 | 82.7% | 2.1DeepSeek Harness Minimal Mode | Public reference · not scored provider-reported | DeepSeek-V4-Flash-0731 official model card and provider evaluation ↗Observed 2026-07-31 · checked 2026-08-24 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 91.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 88.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 88% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 87.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 84.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 83.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 82.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 82% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant SWE-1.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | SWE-1.7 | 81.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 81% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 80.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 80.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 80% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Ornith-1.0-397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-397B | 77.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 76.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 75.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 74.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 70.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna S 2.1 | 70.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 69.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 69.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Composer 2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Composer 2.5 | 69.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 68.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 67.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 66.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 66% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5 | 65.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Ornith-1.0-35B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-35B | 64.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 63.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 63.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Composer 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Composer 2 | 61.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 61.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 60% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 59.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 59.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 59.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 59% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 57% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 56.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 56.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 56.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 56.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 54.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 54% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 52.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 51.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.5 | 50% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 49.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 49.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 47.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 46% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Laguna M.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna M.1 | 45.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Ornith-1.0-9B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-9B | 43.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 41.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 41% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 40.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Laguna XS.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna XS.2 | 35.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 91.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 88.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 88% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 87.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 84.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Terminus-2; as published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | GPT-5.6 Luna | 84.7% | 2.1Terminus 2 harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Fable 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Fable 5 | 84.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 83.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Terminus-2; as published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | Grok 4.5 | 83.3% | 2.1Terminus 2 harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 82.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 82% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant SWE-1.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | SWE-1.7 | 81.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.2 | 81% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 80.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Terminus-2; as published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | Claude Sonnet 5 | 80.4% | 2.1Terminus 2 harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 80.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 80% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| gemini-3.6-flash via Gemini API; default sampling; single-attempt / pass@1 unless noted in Google eval methodology (July 2026). Terminus-2 harness; Terminal-Bench 2.1. source label without registered configuration ID | Gemini 3.6 Flash | 78% | 2.1Terminus 2 harness | Score input provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Ornith-1.0-397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-397B | 77.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Artificial Analysis Terminal-Bench v2.1 independent evaluation (as listed on AA / BenchLM). source label without registered configuration ID | Gemini 3.6 Flash | 77.5% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Gemini 3.6 Flash (Artificial Analysis) ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.3 Codex; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.3-Codex | 77.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 76.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 75.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 74.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Terminus-2; as published in Gemini 3.6 Flash evaluation table for Gemini 3.1 Pro. source label without registered configuration ID | Gemini 3.1 Pro Preview | 73.8% | 2.1Terminus 2 harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 70.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna S 2.1 | 70.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 69.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 69.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Composer 2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Composer 2.5 | 69.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5-Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5-Pro | 68.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-max | DeepSeek V4 Pro | 67.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 66.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M3 | 66% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5 | 65.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.6 | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen 3.6 Max (preview); bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen 3.6 Max (preview) | 65.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Ornith-1.0-35B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-35B | 64.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 63.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5.1 | 63.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-pro-high | DeepSeek V4 Pro | 63.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Composer 2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Composer 2 | 61.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 61.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 mini; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 mini | 60% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Step 3.7 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Step 3.7 Flash | 59.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 59.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 59.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Pro; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Pro | 59.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 59% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiniMax M2.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiniMax M2.7 | 57% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 56.9% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (High); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-high | DeepSeek V4 Flash | 56.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Ultra | 56.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-5 | 56.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Hy3 Preview; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Hy3 Preview | 54.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 54% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| gemini-3.5-flash-lite via Gemini API; configuration as stated in the July 2026 Google launch materials. Terminal-Bench 2.1 as published in Google launch blog. source label without registered configuration ID | Gemini 3.5 Flash-Lite | 54% | 2.1Provider evaluation harness | Public reference · not scored provider-reported | Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 52.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 51.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5 (Reasoning); bulk export does not retain a complete upstream harness configuration. kimi-k2-5-thinking | Kimi K2.5 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.5 | 50.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.5 | 50% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 49.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V4 Flash | 49.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 47.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4 nano; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 nano | 46.3% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MAI-Thinking-1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MAI-Thinking-1 | 46% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Laguna M.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna M.1 | 45.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Ornith-1.0-9B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Ornith-1.0-9B | 43.1% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 27B | 41.6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 41% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-35B-A3B | 40.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Laguna XS.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna XS.2 | 35.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| As published in Gemini 3.5 Flash-Lite launch materials (prior Flash-Lite comparison). source label without registered configuration ID | Gemini 3.1 Flash-Lite | 31% | 2.1Provider evaluation harness | Public reference · not scored provider-reported | Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public terminal-bench leaderboard page (verified 2026-07-21). source label without registered configuration ID | Claude Opus 4.7 | 69.4% | 2.0/2.1BenchLM aggregated public evaluation | Score input source-checked | Terminal-Bench 2.0 Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public terminal-bench leaderboard page (verified 2026-07-21). source label without registered configuration ID | Qwen3.6 Max | 65.4% | 2.0/2.1BenchLM aggregated public evaluation | Score input source-checked | Terminal-Bench 2.0 Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public terminal-bench leaderboard page (verified 2026-07-21). source label without registered configuration ID | Hy3 | 54.4% | 2.0/2.1BenchLM aggregated public evaluation | Score input source-checked | Terminal-Bench 2.0 Leaderboard & Scores — July 2026 ↗Observed 2026-07-20 · checked 2026-07-21 |
|---|
| KimiCode harness; reasoning max; Terminal-Bench 2.1 as published in the Kimi K3 launch blog. kimi-k3-max | Kimi K3 | 88.3% | 2.1KimiCode | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Muse Spark 1.1 (xhigh); Artificial Analysis independent evaluation configuration as published. muse-spark-1-1-xhigh | Muse Spark 1.1 | 77.9% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Muse Spark 1.1 (xhigh) analysis ↗Observed 2026-07-16 · checked 2026-07-16 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Sol | 91.9% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Sol | 88.0% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-sol ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Terra | 88.0% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-terra ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Mythos 5 | 88% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Terra | 87.4% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Luna | 84.7% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Fable 5 | 84.6% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-fable-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Opus 4.8 | 84.6% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-opus-4-8 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Fable 5 | 84.3% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.5 | 84.3% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Grok 4.5 | 83.3% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 82% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.5 | 81.7% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GLM-5.2 | 81% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.6 Luna | 80.9% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-6-luna ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Sonnet 5 | 80.5% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for claude-sonnet-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 5 | 80.4% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark 1.1 | 80% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | GPT-5.4 | 78.3% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gpt-5-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 76.2% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.8 | 74.6% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Qwen3.7-Max | 74.5% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for qwen3-7-max ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 73.8% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for gemini-3-1-pro-preview ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 70.3% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Max | 69.7% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | MiMo-V2.5-Pro | 68.4% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.7 Code source label without registered configuration ID | Kimi K2.7 Code | 67.4% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | MiniMax M3 | 66% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2.6 source label without registered configuration ID | Kimi K2.6 | 65.9% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | MiniMax M3 | 65.2% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for minimax-m3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Pro | 64.0% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-pro ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiMo-V2.5 source label without registered configuration ID | MiMo-V2.5 | 63.7% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | DeepSeek V4 Flash | 61.8% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for deepseek-v4-flash ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6 27B (Reasoning) source label without registered configuration ID | Qwen3.6 27B | 60.7% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.4 mini (xhigh) gpt-5-4-mini-xhigh | GPT-5.4 mini | 59.2% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 59.1% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 55.8% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nemotron 3 Ultra 550B A55B (Reasoning) source label without registered configuration ID | Nemotron 3 Ultra | 53.9% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 397B A17B (Reasoning) source label without registered configuration ID | Qwen3.5 397B A17B | 51.3% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Medium 3.5 | 50.6% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-medium-3-5 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 49.4% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 49.1% | 2.1BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.5 122B A10B (Reasoning) source label without registered configuration ID | Qwen3.5 122B | 47.6% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 46.8% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 45.3% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 44.9% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Grok 4.3 | 39.7% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for grok-4-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemma 4 26B A4B (Reasoning) source label without registered configuration ID | Gemma 4 26B | 39.0% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| NVIDIA Nemotron 3 Super 120B A12B (Reasoning) source label without registered configuration ID | Nemotron 3 Super | 38.6% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 36.3% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 35.2% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 29.6% | 2.1Artificial Analysis independent evaluation | Score input source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 28.5% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Command A+ source label without registered configuration ID | Command A+ | 22.8% | 2.1Artificial Analysis independent evaluation | Public reference · not scored source-checked | Terminal-Bench Hard Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Small 4 | 21.0% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-small-4 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Large 3 | 12.0% | 2.1Artificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-large-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Codex; reasoning max; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | GPT-5.6 Terra | 78.4% | 2.1Codex | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-11 · checked 2026-08-16 |
|---|
| Codex; reasoning max; Terminal-Bench 2.1 verified submission. gpt-5-6-terra-max | GPT-5.6 Terra | 78.4% | 2.1Codex | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-11 · checked 2026-07-15 |
|---|
| Codex; reasoning max; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | GPT-5.6 Luna | 75.7% | 2.1Codex | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-11 · checked 2026-08-16 |
|---|
| Codex; reasoning max; Terminal-Bench 2.1 verified submission. gpt-5-6-luna-max | GPT-5.6 Luna | 75.7% | 2.1Codex | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-11 · checked 2026-07-15 |
|---|
| Muse Spark 1.1; provider evaluation configuration as published in the Meta Muse Spark 1.1 evaluation report (Terminal-Bench 2.0). muse-spark-1-1-xhigh | Muse Spark 1.1 | 80% | 2.1Provider evaluation harness | Public reference · not scored provider-reported | Meta AI Muse Spark 1.1 evaluation report ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|
| Cursor CLI; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Grok 4.5 | 79.3% | 2.1Cursor CLI | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-08-16 |
|---|
| Claude Code; reasoning high; Terminal-Bench 2.1 verified submission. claude-opus-4-8-high | Claude Opus 4.8 | 78.9% | 2.1Claude Code | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-07-15 |
|---|
| Claude Code; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Claude Opus 4.8 | 78.9% | 2.1Claude Code | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-08-16 |
|---|
| mini-SWE-agent; reasoning xhigh; Terminal-Bench 2.1 verified submission. muse-spark-1-1-xhigh | Muse Spark 1.1 | 76.2% | 2.1mini-SWE-agent | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|
| mini-SWE-agent; reasoning xhigh; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Muse Spark 1.1 | 76.2% | 2.1mini-SWE-agent | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-08-16 |
|---|
| Claude Code; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Claude Sonnet 5 | 74.6% | 2.1Claude Code | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-08-16 |
|---|
| Claude Code; reasoning high; Terminal-Bench 2.1 verified submission. claude-sonnet-5-high | Claude Sonnet 5 | 74.6% | 2.1Claude Code | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-07-09 · checked 2026-07-15 |
|---|
| LongCat-2.0 (in-house harness) longcat-2-0-default | LongCat-2.0 | 70.8% | 2.1LongCat unified in-house harness | Score input provider-reported | LongCat-2.0 model card ↗Observed 2026-06-30 · checked 2026-08-15 |
|---|
| Claude Code; reasoning xhigh; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Claude Fable 5 | 83.8% | 2.1Claude Code | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-06-07 · checked 2026-08-16 |
|---|
| Claude Code; reasoning xhigh; Terminal-Bench 2.1 verified submission. claude-fable-5-xhigh | Claude Fable 5 | 83.8% | 2.1Claude Code | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-06-07 · checked 2026-07-15 |
|---|
| Terminus 2; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Claude Fable 5 | 80.5% | 2.1Terminus 2 | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-06-05 · checked 2026-08-16 |
|---|
| Terminus 2; reasoning high; Terminal-Bench 2.1 verified submission. claude-fable-5-high | Claude Fable 5 | 80.4% | 2.1Terminus 2 | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-06-05 · checked 2026-07-15 |
|---|
| Terminus-2 agent system; provider-published Terminal-Bench 2.1 evaluation. gpt-5-5-high | GPT-5.5 | 78.2% | 2.1Terminus 2 | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Terminus-2 agent system; provider-published Terminal-Bench 2.1 evaluation. gemini-3-5-flash-high | Gemini 3.5 Flash | 76.2% | 2.1Terminus 2 | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Terminus-2 agent system; provider-published Terminal-Bench 2.1 evaluation. source label without registered configuration ID | Gemini 3.1 Pro Preview | 70.3% | 2.1Terminus 2 | Source carrier · no additional scoring weight provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Codex; reasoning xhigh; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | GPT-5.5 | 83.2% | 2.1Codex | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-08-16 |
|---|
| Codex; reasoning xhigh; Terminal-Bench 2.1 verified submission. gpt-5-5-xhigh | GPT-5.5 | 83.1% | 2.1Codex | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-07-15 |
|---|
| Terminus 2; reasoning xhigh; Terminal-Bench 2.1 verified submission. gpt-5-5-xhigh | GPT-5.5 | 78% | 2.1Terminus 2 | Public reference · not scored official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-07-15 |
|---|
| Terminus 2; reasoning xhigh; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | GPT-5.5 | 78.0% | 2.1Terminus 2 | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-08-16 |
|---|
| Opus 4.7; reasoning=max; agent=Claude Code claude-opus-4-7-max | Claude Opus 4.7 | 68.9% | 2.1Claude Code | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-08-16 |
|---|
| Opus 4.7; reasoning=max; agent=Terminus 2 claude-opus-4-7-max | Claude Opus 4.7 | 66.1% | 2.1Terminus 2 | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-08-16 |
|---|
| Gemini CLI; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Gemini 3.1 Pro Preview | 65.8% | 2.1Gemini CLI | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-07-15 |
|---|
| Terminus 2; reasoning high; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | Gemini 3.1 Pro Preview | 65.6% | 2.1Terminus 2 | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-07-15 |
|---|
| Claude Code; reasoning max; Terminal-Bench 2.1 verified submission. source label without registered configuration ID | GLM-5.1 | 58.7% | 2.1Claude Code | Score input official-leaderboard | Terminal-Bench 2.1 leaderboard ↗Observed 2026-05-01 · checked 2026-07-16 |
|---|