| DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 54.4% | 1.1mini-swe-agent | Source carrier · no additional scoring weight provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 42.2% | 1.1mini-swe-agent | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 16.5% | 1.1mini-swe-agent | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0 deepseek-v4-flash-vision-exp-max-harness | DeepSeek V4 Flash Vision Exp | 59.3% | v1.1DeepSeek Harness Minimal Mode | Public reference · not scored provider-reported | DeepSeek-V4-Flash-Vision-Exp release and provider evaluation ↗Observed 2026-08-21 · checked 2026-08-24 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 42.2% | 1.1Qwen3.8-27B official text table | Public reference · not scored provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 14.2% | 1.1Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 13.3% | 1.1Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| GPT-5.6 Sol as published by Z.AI gpt-5-6-sol-max | GPT-5.6 Sol | 72.7% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Fable 5 w/ fallback as published by Z.AI claude-fable-5-max | Claude Fable 5 | 69.7% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Kimi K3 as published by Z.AI kimi-k3-max | Kimi K3 | 67.5% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 66.9% | 1.1Z.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| DeepSeek V4 Pro 0813 as published by Z.AI deepseek-v4-pro-0813-max | DeepSeek V4 Pro 0813 | 62.7% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Opus 4.8 as published by Z.AI claude-opus-4-8-max | Claude Opus 4.8 | 58% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Qwen3.8 Max as published by Z.AI qwen-3-8-max-xhigh | Qwen3.8 Max | 56.6% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.2 as published by Z.AI glm-5-2-max | GLM-5.2 | 46.2% | 1.1Z.AI GLM-5.3 official comparison table | Relative comparison · not standard direct provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Fable 5 (with fallback) claude-fable-5-deepseek-0813-release-unspecified | Claude Fable 5 | 70% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GPT-5.6 Terra gpt-5-6-terra-max | GPT-5.6 Terra | 69.6% | v1.1DeepSWE v1.1 public leaderboard; Gemini 3.7 self-run used mini SWE agent and LiteLLM 1.96 | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Kimi K3 kimi-k3-deepseek-0813-release-unspecified | Kimi K3 | 67.5% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.7 Flash gemini-3-7-flash-high | Gemini 3.7 Flash | 65.3% | v1.1DeepSWE v1.1 public leaderboard; Gemini 3.7 self-run used mini SWE agent and LiteLLM 1.96 | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Pro 0813 deepseek-v4-pro-0813-deepseek-0813-release-unspecified | DeepSeek V4 Pro 0813 | 62.7% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Opus 4.8 claude-opus-4-8-deepseek-0813-release-unspecified | Claude Opus 4.8 | 58% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Muse Spark 1.2 muse-spark-1-2-xhigh | Muse Spark 1.2 | 54.9% | v1.1DeepSWE v1.1 public leaderboard; Gemini 3.7 self-run used mini SWE agent and LiteLLM 1.96 | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Flash 0731 deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 54.4% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Source carrier · no additional scoring weight provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Sonnet 5 claude-sonnet-5-max | Claude Sonnet 5 | 53.8% | v1.1DeepSWE v1.1 public leaderboard; Gemini 3.7 self-run used mini SWE agent and LiteLLM 1.96 | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.6 Flash gemini-3-6-flash-high | Gemini 3.6 Flash | 48.6% | v1.1DeepSWE v1.1 public leaderboard; Gemini 3.7 self-run used mini SWE agent and LiteLLM 1.96 | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GLM-5.2 glm-5-2-deepseek-0813-release-unspecified | GLM-5.2 | 46.2% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Pro Preview deepseek-v4-pro-deepseek-0813-release-unspecified | DeepSeek V4 Pro | 12.8% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| DeepSeek V4 Flash Preview deepseek-v4-flash-deepseek-0813-release-unspecified | DeepSeek V4 Flash | 7.3% | v1.1DeepSeek V4 Pro 0813 official GA comparison table | Public reference · not scored provider-reported | DeepSeek permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GPT-5.6 Sol Max gpt-5-6-sol-max | GPT-5.6 Sol | 73% | v1.1xAI Grok 4.6 official release comparison table | Public reference · not scored provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Fable 5 Max claude-fable-5-max | Claude Fable 5 | 70% | v1.1xAI Grok 4.6 official release comparison table | Public reference · not scored provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Grok 4.6 High grok-4-6-high | Grok 4.6 | 65.9% | v1.1xAI Grok 4.6 official release comparison table | Public reference · not scored provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Grok 4.5 High grok-4-5-aa-2-high | Grok 4.5 | 54% | v1.1xAI Grok 4.6 official release comparison table | Public reference · not scored provider-reported | Grok 4.6 ↗Observed 2026-08-12 · checked 2026-08-12 |
|---|
| Claude Opus 5 (reasoning configuration not stated in chart) claude-opus-5-meta-muse-12-release-unspecified | Claude Opus 5 | 65% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| GPT-5.6 Terra (reasoning configuration not stated in chart) gpt-5-6-terra-meta-muse-12-release-unspecified | GPT-5.6 Terra | 64.8% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.2 (xhigh) muse-spark-1-2-xhigh | Muse Spark 1.2 | 59.3% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Public reference · not scored provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.2 (xhigh) with Muse Code muse-spark-1-2-xhigh | Muse Spark 1.2 | 59.3% | 1.1Muse Code evaluation on 113 DeepSWE 1.1 tasks | Public reference · not scored provider-reported | Muse Spark 1.2 Evaluation Methodology ↗Observed 2026-08-05 · checked 2026-08-05 |
|---|
| Grok 4.5 (reasoning configuration not stated in chart) grok-4-5-meta-muse-12-release-unspecified | Grok 4.5 | 56.6% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Muse Spark 1.1 (reasoning configuration not stated in chart) muse-spark-1-1-meta-muse-12-release-unspecified | Muse Spark 1.1 | 53% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Gemini 3.6 Flash (reasoning configuration not stated in chart) gemini-3-6-flash-meta-muse-12-release-unspecified | Gemini 3.6 Flash | 40% | v1.1Meta release chart for DeepSWE 1.1; the chart labels each agent scaffold separately | Relative comparison · not standard direct provider-reported | Meta Superintelligence Labs permanent refresh source ↗Observed 2026-08-05 · checked 2026-08-15 |
|---|
| Qwen3.8 Max (xhigh, default reasoning effort) qwen-3-8-max-xhigh | Qwen3.8 Max | 56.6% | 1.1Best of Claude Code and mini-SWE-agent evaluations | Public reference · not scored provider-reported | Qwen3.8 Max official release ↗Observed 2026-08-03 · checked 2026-08-05 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 72.7 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 69.6 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 68.8 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 67.5 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 67.2 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V4 Flash (Max); bulk export does not retain a complete upstream harness configuration. deepseek-v4-flash-max | DeepSeek V4 Flash | 54.4 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 53.3 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 53 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 49 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna S 2.1 | 40.4 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| DeepSeek Harness Minimal Mode; max effort; top_p=0.95; temperature=1.0 deepseek-v4-flash-0731-max-harness | DeepSeek V4 Flash 0731 | 54.4% | v1.1DeepSeek Harness Minimal Mode | Public reference · not scored provider-reported | DeepSeek-V4-Flash-0731 official model card and provider evaluation ↗Observed 2026-07-31 · checked 2026-08-24 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 72.7 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 69.6 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 5 | 68.8 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 67.5 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 67.2 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 53.3 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 53 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 49 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna S 2.1 | 40.4 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| DeepSWE v1.1 agentic coding as published in the Anthropic Claude Opus 5 launch table. source label without registered configuration ID | Claude Opus 5 | 68.8% | v1.1Anthropic provider evaluation | Public reference · not scored provider-reported | Introducing Claude Opus 5 ↗Observed 2026-07-24 · checked 2026-07-24 |
|---|
| As published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | GPT-5.6 Luna | 67% | v1.1Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| As published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | Grok 4.5 | 54% | v1.1Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| As published in Gemini 3.6 Flash evaluation table. source label without registered configuration ID | Claude Sonnet 5 | 54% | v1.1Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| gemini-3.6-flash via Gemini API; default sampling; single-attempt / pass@1 unless noted in Google eval methodology (July 2026). High reasoning; public DeepSWE v1.1 leaderboard as cited by Google. gemini-3-6-flash-high | Gemini 3.6 Flash | 49% | v1.1Provider evaluation harness | Public reference · not scored provider-reported | Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| As published in Gemini 3.6 Flash evaluation table for Gemini 3.5 Flash comparison. source label without registered configuration ID | Gemini 3.5 Flash | 37% | v1.1Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 72.7 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 69.6 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 67.5 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 67.2 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 53.3 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.5 | 53 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.6 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.6 Flash | 49 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Laguna S 2.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Laguna S 2.1 | 40.4 USD | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| mini-SWE-agent common harness; DeepSWE v1.1; max reasoning. Provider also reports 67.5 with KimiCode. kimi-k3-max | Kimi K3 | 67.3% | v1.1mini-SWE-agent | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|