| GPT-5.6 Sol as published by Z.AI gpt-5-6-sol-max | GPT-5.6 Sol | 293 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Fable 5 w/ fallback as published by Z.AI claude-fable-5-max | Claude Fable 5 | 247 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GPT-5.6 Sol as published by Z.AI gpt-5-6-sol-max | GPT-5.6 Sol | 216 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Fable 5 w/ fallback as published by Z.AI claude-fable-5-max | Claude Fable 5 | 181 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 130 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Opus 4.8 as published by Z.AI claude-opus-4-8-max | Claude Opus 4.8 | 120 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.3 (max) glm-5-3-max | GLM-5.3 | 105 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Claude Opus 4.8 as published by Z.AI claude-opus-4-8-max | Claude Opus 4.8 | 80 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Kimi K3 as published by Z.AI kimi-k3-max | Kimi K3 | 70 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.2 as published by Z.AI glm-5-2-max | GLM-5.2 | 39 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Kimi K3 as published by Z.AI kimi-k3-max | Kimi K3 | 36 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| GLM-5.2 as published by Z.AI glm-5-2-max | GLM-5.2 | 29 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Qwen3.8 Max as published by Z.AI qwen-3-8-max-xhigh | Qwen3.8 Max | 26 tasks | 6hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Qwen3.8 Max as published by Z.AI qwen-3-8-max-xhigh | Qwen3.8 Max | 14 tasks | 2hZ.AI GLM-5.3 official comparison table | Public reference · not scored provider-reported | GLM-5.3: Frontier Coding with Emergent Cyber Capabilities ↗Observed 2026-08-14 · checked 2026-08-15 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 23.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 13.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 12.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 0.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-08-01 ↗Observed 2026-08-01 · checked 2026-08-01 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 23.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 13.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 12.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 0.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27 |
|---|
| Exact BenchLM registry variant GPT-5.6 Sol; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Sol | 33.7% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Terra; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Terra | 23.2% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Claude Mythos 5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Claude Mythos 5 | 17.5% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.5; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.5 | 13.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.6 Luna; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.6 Luna | 12.4% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant GPT-5.4; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | GPT-5.4 | 6% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact BenchLM registry variant Muse Spark 1.1; bulk export does not retain a complete upstream harness configuration. source label without registered configuration ID | Muse Spark 1.1 | 0.8% | 2026Source-native system | Source carrier · no additional scoring weight source-checked | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Sol | 33.7% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Terra | 23.2% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Claude Mythos 5 | 17.5% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.5 | 13.4% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.6 Luna | 12.4% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | GPT-5.4 | 6% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Exact model variant as listed on the BenchLM public benchmark leaderboard page. source label without registered configuration ID | Muse Spark 1.1 | 0.8% | 2026-05BenchLM aggregated public evaluation | Source carrier · no additional scoring weight source-checked | BenchLM public benchmark leaderboards ↗Observed 2026-07-15 · checked 2026-07-15 |
|---|
| Muse Spark 1.1; provider evaluation configuration as published in the Meta Muse Spark 1.1 evaluation report (ExploitGym). muse-spark-1-1-xhigh | Muse Spark 1.1 | 0.8% | 2026-05Provider evaluation harness | Public reference · not scored provider-reported | Meta AI Muse Spark 1.1 evaluation report ↗Observed 2026-07-09 · checked 2026-07-16 |
|---|