GLM-5.2 vs GLM-5.3
6 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
6| Benchmark | GLM-5.2 | GLM-5.3 |
|---|---|---|
AutomationBench 1.0.6Potentially non-equivalent: Evaluator sets differ. | 26.2%1.0.6 · Task pass rate Zapier AutomationBench Team | 48.2%1.0.6 · Task pass rate Z.ai |
Potentially non-equivalent · Evaluation details for AutomationBench 1.0.6Evaluator sets differ. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||
77.2% | 84.5% | |
Evaluation details for CyberGym OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||
44.0% | 69.0% | |
Evaluation details for DeepSWE 1.1The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||
77.8% | 84.2% | |
Evaluation details for MCP Atlas Public, April 2026 updateThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||
48.9% | 58.0% | |
Evaluation details for NL2Repo Bench OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 81.0%2.1 — Terminus-2 · Terminal-Bench Accuracy Z.ai | 88.2%2.1 — Claude Code · Terminal-Bench Accuracy Z.ai |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. GLM-5.2 GLM-5.3 | ||