| Qwen3.8-Flash-Next (xhigh default thinking configuration) qwen-3-8-flash-next-xhigh | Qwen3.8-Flash-Next | 90.6% | RQ With CIsystem:qwen3-8-flash-next:charxiv-with-ci | Score input source-checked | Qwen3.8-Flash-Next current official launch page ↗Observed 2026-08-26 · checked 2026-08-29 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 90.2% | RQ With CIsystem:qwen3-8-flash-next-comparison:cell:vision:charxiv-with-ci:1 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| GLM-5.3-Flash (max effort, documented default) glm-5-3-flash-max | GLM-5.3-Flash | 89.4% | May 2026system:glm-5-3-flash-launch:charxiv:0 | Score input source-checked | GLM-5.3-Flash individual evaluations ↗Observed 2026-08-26 · checked 2026-08-27 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 85.9% | RQ With CIsystem:qwen3-8-flash-next-comparison:cell:vision:charxiv-with-ci:2 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.7-Plus (provider-published configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-7-plus-unspecified | Qwen3.7-Plus | 85.8% | RQ Without CIsystem:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:2 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Qwen3.8-27B (xhigh default thinking configuration) as published in Qwen's Qwen3.8-Flash-Next comparison table qwen-3-8-27b-xhigh | Qwen3.8-27B | 83.7% | RQ Without CIsystem:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:1 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Claude Opus 4.6 (Max) as published in Qwen's Qwen3.8-Flash-Next comparison table claude-opus-4-6-max | Claude Opus 4.6 | 66% | RQ Without CIsystem:qwen3-8-flash-next-comparison:cell:vision:charxiv-no-ci:3 | Relative comparison · not standard direct provider-reported | Qwen3.8-Flash-Next launch and official provider evaluations ↗Observed 2026-08-26 · checked 2026-08-26 |
|---|
| Claude Opus 5 (max) as published in Meta's comparison table claude-opus-5-max | Claude Opus 5 | 89.3% | May 2026Meta evaluation container | Relative comparison · not standard direct provider-reported | Multimodal Intelligence of Muse Spark 1.2 ↗Observed 2026-08-20 · checked 2026-08-27 |
|---|
| Gemini 3.7 Flash (high) as published in Meta's comparison table gemini-3-7-flash-high | Gemini 3.7 Flash | 88.7% | May 2026Meta evaluation container | Relative comparison · not standard direct provider-reported | Multimodal Intelligence of Muse Spark 1.2 ↗Observed 2026-08-20 · checked 2026-08-27 |
|---|
| Muse Spark 1.1 (xhigh) as published in Meta's comparison table muse-spark-1-1-xhigh | Muse Spark 1.1 | 88.4% | May 2026Meta evaluation container | Relative comparison · not standard direct provider-reported | Multimodal Intelligence of Muse Spark 1.2 ↗Observed 2026-08-20 · checked 2026-08-27 |
|---|
| Muse Spark 1.2 (xhigh) muse-spark-1-2-xhigh | Muse Spark 1.2 | 87.6% | May 2026Meta evaluation container | Score input source-checked | Muse Spark 1.2 individual evaluations ↗Observed 2026-08-20 · checked 2026-08-27 |
|---|
| GPT-5.6 Sol (max) as published in Meta's comparison table gpt-5-6-sol-max | GPT-5.6 Sol | 84% | May 2026Meta evaluation container | Relative comparison · not standard direct provider-reported | Multimodal Intelligence of Muse Spark 1.2 ↗Observed 2026-08-20 · checked 2026-08-27 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 90.2% | RQ With CIQwen3.8-27B official VL table | Public reference · not scored provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 85.9% | RQ With CIQwen3.8-27B official VL table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.7-Plus as published by Qwen qwen-3-7-plus-unspecified | Qwen3.7-Plus | 85.8% | RQ Without CIQwen3.8-27B official VL table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.8-27B (xhigh) qwen-3-8-27b-xhigh | Qwen3.8-27B | 83.7% | RQ Without CIQwen3.8-27B official VL table | Score input provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Qwen3.6-27B as published by Qwen qwen3-6-27b-default | Qwen3.6 27B | 78.4% | RQ Without CIQwen3.8-27B official VL table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Claude Opus 4.6 Max as published by Qwen claude-opus-4-6-max | Claude Opus 4.6 | 66% | RQ Without CIQwen3.8-27B official VL table | Relative comparison · not standard direct provider-reported | Qwen3.8-27B official model card ↗Observed 2026-08-15 · checked 2026-08-15 |
|---|
| Gemini 3.6 Flash gemini-3-6-flash-high | Gemini 3.6 Flash | 89.4% | May 2026CharXiv Reasoning with tools | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.7 Flash gemini-3-7-flash-medium | Gemini 3.7 Flash | 88.7% | May 2026CharXiv Reasoning with tools | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Sonnet 5 claude-sonnet-5-max | Claude Sonnet 5 | 88.3% | May 2026CharXiv Reasoning with tools | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| GPT-5.6 Terra gpt-5-6-terra-max | GPT-5.6 Terra | 85.9% | May 2026CharXiv Reasoning without tools; Gemini and GPT values self-computed | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.6 Flash gemini-3-6-flash-high | Gemini 3.6 Flash | 85.2% | May 2026CharXiv Reasoning without tools; Gemini and GPT values self-computed | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Gemini 3.7 Flash gemini-3-7-flash-medium | Gemini 3.7 Flash | 84.5% | May 2026CharXiv Reasoning without tools; Gemini and GPT values self-computed | Score input provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Claude Sonnet 5 claude-sonnet-5-max | Claude Sonnet 5 | 77% | May 2026CharXiv Reasoning without tools; Gemini and GPT values self-computed | Public reference · not scored provider-reported | Google DeepMind permanent refresh source ↗Observed 2026-08-13 · checked 2026-08-13 |
|---|
| Muse Glimmer-30B; high reasoning; temperature=1.0; top_p=0.95; top_k=64 muse-glimmer-30b-high | Muse Glimmer 30B | 78.8% | Reasoning / validation setMeta evaluation of 1,000 CharXiv validation questions, four attempts | Public reference · not scored provider-reported | Muse Glimmer Evaluation Methodology ↗Observed 2026-08-10 · checked 2026-08-10 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 93.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Kimi K3 public default configuration (reasoning_effort=max; temperature=1.0) kimi-k3-max | Kimi K3 | 91.3% | CharXiv (RQ) with PythonMoonshot Kimi K3 model-card evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-29 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 91% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 89.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 88.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 88.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 86.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 86.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 85.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 85.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 82.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 82.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 82% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 81.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Inkling-Small; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling-Small | 81.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5 | 81% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 80.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 78.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 78% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 77.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 77.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 76.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 73.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 68.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 60.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 52.7% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 93.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 91.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 91% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 89.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 88.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 88.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 86.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 86.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 85.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 85.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 82.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 82.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 82% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 81.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5 | 81% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 80.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 78.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 78% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 77.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 77.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 76.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 73.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 68.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 60.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 52.7% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 93.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K3; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K3 | 91.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.7 (Adaptive); bulk export does not retain a complete upstream harness configuration. claude-opus-4-7-max | Claude Opus 4.7 | 91% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.8; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.8 | 89.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| gemini-3.6-flash via Gemini API; default sampling; single-attempt / pass@1 unless noted in Google eval methodology (July 2026). CharXiv Reasoning with search and code-execution tools. source label without registered configuration ID | Gemini 3.6 Flash | 89.4% | May 2026search+code-execution harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 88.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 5 | 88.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sakana Fugu-Ultra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu-Ultra | 86.6% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark | 86.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.7 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.7-Plus | 85.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| gemini-3.6-flash via Gemini API; default sampling; single-attempt / pass@1 unless noted in Google eval methodology (July 2026). CharXiv Reasoning without tools; self-computed. source label without registered configuration ID | Gemini 3.6 Flash | 85.2% | May 2026Provider evaluation harness | Score input provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Sakana Fugu; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Sakana Fugu | 85.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.5 Flash; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.5 Flash | 84.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 82.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| CharXiv Reasoning no-tools; as published in Gemini 3.6 evaluation table. source label without registered configuration ID | GPT-5.6 Luna | 82.7% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.2; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.2 | 82.1% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Inkling; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Inkling | 82% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| CharXiv Reasoning no-tools; as published in Gemini 3.6 evaluation table. source label without registered configuration ID | Grok 4.5 | 81.6% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6 Plus; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 Plus | 81.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant MiMo-V2.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | MiMo-V2.5 | 81% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5 397B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5 397B A17B | 80.8% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Kimi K2.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Kimi K2.6 | 80.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-27B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6 27B | 78.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.6-35B-A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.6-35B-A3B | 78% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Sonnet 4.6; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Sonnet 4.6 | 77.4% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Qwen3.5-122B-A10B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Qwen3.5-122B-A10B | 77.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| CharXiv Reasoning no-tools; as published in Gemini 3.6 evaluation table. source label without registered configuration ID | Claude Sonnet 5 | 77% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.6 Flash product page with evaluation table ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Nemotron 3 Nano Omni 30B A3B; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Nemotron 3 Nano Omni 30B A3B | 76.3% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Gemini 3.1 Flash-Lite; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 73.2% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Opus 4.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Opus 4.5 | 68.5% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Grok 4.20; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Grok 4.20 | 60.9% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Command A+; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Command A+ | 52.7% | 2024Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| CharXiv Reasoning with Python; max reasoning; average of three runs. kimi-k3-max | Kimi K3 | 91.3% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Kimi K3: Open Frontier Intelligence ↗Observed 2026-07-16 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Mythos 5 | 93.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.8 | 89.9% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark 1.1 | 88.4% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 5 | 88.3% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark | 86.4% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.7-Plus | 85.9% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.5 Flash | 84.2% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 82.8% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Qwen3.6 Plus | 81.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.1 Pro Preview | 80.2% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Sonnet 4.6 | 77.4% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Gemini 3.1 Flash-Lite | 73.2% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Opus 4.5 | 68.5% | May 2026BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Muse Spark 1.1; provider evaluation configuration as published in the Meta Muse Spark 1.1 evaluation report (CharXiv Reasoning). muse-spark-1-1-xhigh | Muse Spark 1.1 | 88.4% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Meta AI Muse Spark 1.1 evaluation report ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.5 Flash | 84.2% | May 2026Provider evaluation harness | Score input provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | GPT-5.5 | 84.1% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|
| Exact model variant and provider evaluation configuration stated in the Gemini 3.5 Flash model card; single-attempt where specified. source label without registered configuration ID | Gemini 3.1 Pro Preview | 83.3% | May 2026Provider evaluation harness | Public reference · not scored provider-reported | Gemini 3.5 Flash model card ↗Observed 2026-05-19 · checked 2026-07-15 |
|---|