GLM-5.1 vs GLM-5.2
6 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
6| Benchmark | GLM-5.1 | GLM-5.2 |
|---|---|---|
68.7% | 77.2% | |
Evaluation details for CyberGym OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||
86.2% | 91.2% | |
Evaluation details for GPQA DiamondThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||
75.6% | 77.8% | |
Evaluation details for MCP Atlas Public, April 2026 updateThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||
42.7% | 48.9% | |
Evaluation details for NL2Repo Bench OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||
58.4% | 62.1% | |
Evaluation details for SWE-bench Pro PublicThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 63.5%2.0 — Terminus-2 · Terminal-Bench Accuracy Z.ai | 81.0%2.1 — Terminus-2 · Terminal-Bench Accuracy Z.ai |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. GLM-5.1 GLM-5.2 | ||