| AA Long Context Reasoningstandard | 76.667% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis AA-LCR evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| CritPtstandard | 20.857% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis CritPt evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| GDPval-AA v2v2 | 49.878% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis GDPval-AA v2 | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| GPQA Diamonddiamond | 89.495% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis GPQA Diamond evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| Humanity's Last Examtext-only current | 41.149% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis text-only HLE evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| SciCode2024 | 50.463% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis SciCode evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · reference-only |
|---|
| τ³-Banking3 | 34.639% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis tau3-Banking evaluation | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| Terminal-Bench2.1 | 77.903% | GLM-5.2 (max) (Artificial Analysis independent run)Artificial Analysis Terminal-Bench v2.1 in e2b | GLM-5.2 (max) individual evaluations ↗Observed 2026-08-29 · checked 2026-08-29independently-verified · ranking-eligible |
|---|
| Artificial Analysis Agentic Index2026 | 43.06 index | Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 21 July 2026 ↗Observed 2026-07-21 · checked 2026-07-21source-checked · reference-only |
|---|
| Artificial Analysis Agentic Index2026 | 43.06 index | Exact BenchLM registry variant GLM-5.2; bulk export does not retain a complete upstream harness configuration.Source-native system | BenchLM public datasets — 2026-07-27 ↗Observed 2026-07-27 · checked 2026-07-27source-checked · reference-only |
|---|