GLM-5.1 · GLM-5.2

Shared benchmarks

6
Shared benchmarks for GLM-5.1 and GLM-5.2
BenchmarkGLM-5.1GLM-5.2
68.7%
77.2%
Evaluation details for CyberGym Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2

86.2%
91.2%
Evaluation details for GPQA Diamond

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2

75.6%
77.8%
Evaluation details for MCP Atlas Public, April 2026 update

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2

42.7%
48.9%
Evaluation details for NL2Repo Bench Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2

58.4%
62.1%
Evaluation details for SWE-bench Pro Public

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2

Terminal-BenchPotentially non-equivalent: Benchmark versions differ.
63.5%2.0 — Terminus-2 · Terminal-Bench Accuracy
Z.ai
81.0%2.1 — Terminus-2 · Terminal-Bench Accuracy
Z.ai
Potentially non-equivalent · Evaluation details for Terminal-Bench

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

GLM-5.1

GLM-5.2