| Qwen3.8-Flash-Next (xhigh default thinking configuration) qwen-3-8-flash-next-xhigh | Qwen3.8-Flash-Next | 91.9% | v6system:qwen3-8-flash-next:lcb-v6 | Score input source-checked | Qwen3.8-Flash-Next current official launch page ↗Observed 2026-08-26 · checked 2026-08-29 |
|---|
| DeepSeek-V4-Flash-0731 (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table deepseek-v4-flash-0731-deepseek-0813-release-unspecified | DeepSeek V4 Flash 0731 | 90.6% | v6system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:3 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 90.3% | v6system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:1 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 89.6% | v6system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:2 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison table claude-opus-4-6-max | Claude Opus 4.6 | 88.8% | v6system:qwen3-8-flash-next-comparison:cell:language:lcb-v6:4 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Solar Open 2 250B (high) solar-open2-250b-high | Solar Open 2 250B | 92.4% | v6Solar Open 2 card English table | Score input provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| DeepSeek V4 Flash max as published by Upstage deepseek-v4-flash-max | DeepSeek V4 Flash | 92.3% | v6Solar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 90.3% | v6Qwen3.8-27B official text table | Score input provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 89.6% | v6Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| MiMo-V2.5 as published by Upstage mimo-v2-5-solar-open2-unspecified | MiMo-V2.5 | 89.1% | v6Solar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Claude Opus 4.6 Max as published by Qwen claude-opus-4-6-max | Claude Opus 4.6 | 88.8% | v6Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Command A+ as published by Upstage command-a-plus-solar-open2-unspecified | Command A+ | 86.1% | v6Solar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Mistral Medium 3.5 as published by Upstage mistral-medium-3-5-high | Mistral Medium 3.5 | 84.9% | v6Solar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 83.9% | v6Qwen3.8-27B official text table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Solar Open 100B as published by Upstage solar-open-100b-reasoning-high | Solar Open 100B (Reasoning) | 56.5% | v6Solar Open 2 card English table | Relative comparison · not standard direct provider-reported | Solar Open 2 250B model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 91.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 89.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 84.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 83.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 37.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 91.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 89.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 84.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 83.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 37.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Max; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Max | 91.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 89.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GLM-4.7; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GLM-4.7 | 84.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 83.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant DeepSeek V3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | DeepSeek V3 | 37.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Max | 91.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 89.6% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.7 (Reasoning) source label without registered configuration ID | GLM-4.7 | 89.4% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5.2 (xhigh) gpt-5-2-xhigh | GPT-5.2 | 88.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 87.1% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.2 (Reasoning) source label without registered configuration ID | DeepSeek V3.2 | 86.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o4-mini (high) o4-mini-high | o4-mini | 85.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 Thinking source label without registered configuration ID | Kimi K2 Thinking | 85.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Kimi K2.5 | 85% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 84.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 (high) gpt-5-high | GPT-5 | 84.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 Codex (high) gpt-5-codex-high | GPT-5 Codex | 84.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 Fast (Reasoning) source label without registered configuration ID | Grok 4 Fast | 83.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2 source label without registered configuration ID | MiniMax M2 | 82.6% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 4 source label without registered configuration ID | Grok 4 | 81.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| MiniMax-M2.1 source label without registered configuration ID | MiniMax M2.1 | 81.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o3 source label without registered configuration ID | o3 | 80.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Pro source label without registered configuration ID | Gemini 2.5 Pro | 80.1% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| DeepSeek V3.1 Terminus (Reasoning) source label without registered configuration ID | DeepSeek V3.1 Terminus | 79.8% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Nova 2.0 Pro Preview (medium) source label without registered configuration ID | Nova 2 Pro | 73.0% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.5 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4.5 | 71.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Gemini 2.5 Flash Preview (Sep '25) (Reasoning) source label without registered configuration ID | Gemini 2.5 Flash | 71.3% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok 3 mini Reasoning (high) source label without registered configuration ID | Grok 3 mini | 69.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GLM-4.6 (Reasoning) source label without registered configuration ID | GLM-4.6 | 69.5% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| GPT-5 mini (medium) gpt-5-mini-medium | GPT-5 mini | 69.2% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| o1 source label without registered configuration ID | o1 | 67.9% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Grok Code Fast 1 source label without registered configuration ID | Grok Code Fast 1 | 65.7% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 4 | 65.5% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4.1 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4.1 | 65.4% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 4 Opus (Reasoning) source label without registered configuration ID | Claude Opus 4 | 63.6% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Kimi K2 0905 source label without registered configuration ID | Kimi K2 0905 | 61.0% | CurrentArtificial Analysis independent evaluation | Public reference · not scored source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Pro | 56.8% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | DeepSeek V4 Flash | 55.2% | CurrentBenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Claude Haiku 4.5 | 51.1% | CurrentArtificial Analysis evaluation harness | Score input source-checked | Artificial Analysis evaluations for claude-4-5-haiku ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Claude 3.7 Sonnet (Reasoning) source label without registered configuration ID | Claude Sonnet 3.7 | 47.3% | CurrentArtificial Analysis independent evaluation | Score input source-checked | Artificial Analysis Models Leaderboard ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Artificial Analysis public model evaluation page; exact provider variant as listed on the page. source label without registered configuration ID | Mistral Large 3 | 46.5% | CurrentArtificial Analysis evaluation harness | Public reference · not scored source-checked | Artificial Analysis evaluations for mistral-large-3 ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Qwen3.6-35B-A3B default thinking configuration qwen3-6-35b-a3b-aa-reasoning-default | Qwen3.6-35B-A3B | 80.4% | v6Qwen3.6 official model-card evaluation | Source carrier · no additional scoring weight source-checked | Qwen3.6-35B-A3B official model card ↗Observed 2026-04-15 · checked 2026-08-29 |
|---|
| DeepSeek-R1-0528 source label without registered configuration ID | DeepSeek R1 0528 (May '25) | 73.1% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| Gemini-2.5-Pro-05-06 source label without registered configuration ID | Gemini 2.5 Pro Preview (May' 25) | 71.8% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| EXAONE-4.0-32B source label without registered configuration ID | Exaone 4.0 32B | 70% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| Claude-Sonnet-4 source label without registered configuration ID | Claude Sonnet 4 | 47.1% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| Claude-Opus-4 source label without registered configuration ID | Claude Opus 4 | 46.9% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| Claude-3.5-Sonnet-20241022 source label without registered configuration ID | Claude 3.5 Sonnet (Oct '24) | 36.4% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| GPT-4O-2024-08-06 source label without registered configuration ID | GPT-4o (Aug '24) | 29.5% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| GPT-4-Turbo-2024-04-09 source label without registered configuration ID | GPT-4 Turbo | 28.7% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| GPT-4O-mini-2024-07-18 source label without registered configuration ID | GPT-4o mini | 27.5% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| DeepSeek-V3 source label without registered configuration ID | DeepSeek V3 | 27.2% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|
| Claude-3-Haiku source label without registered configuration ID | Claude 3 Haiku | 20.2% | standardLiveCodeBench official generation leaderboard | Score input official-leaderboard | LiveCodeBench permanent refresh source ↗Observed 2025-05-01 · checked 2026-08-18 |
|---|