Skip to main content
Benchmark Registry
Models
Benchmarks
Organizations
Search the registry
Search
GPQA
Graduate-Level Google-Proof Q&A
Benchmark
GPQA
Evaluated by
GPQA Authors
Release date
November 20, 2023
Version
GPQA Diamond
Metric
GPQA Diamond Accuracy
Search models
Search
Results
4 results
Rows per page
50
100
500
Apply
Latest
History
All providers
Alibaba (Tongyi)
Anthropic
DeepSeek AI
Google (DeepMind)
Microsoft AI
Mistral AI
Moonshot AI
NVIDIA
OpenAI
Thinking Machines Lab
Z.ai
Graduate-Level Google-Proof Q&A Diamond results
Provider
Model
Score
Source
Registry No.
Z.ai
GLM-5.2
91.2%
Source
for GLM-5.2 on Graduate-Level Google-Proof Q&A Diamond
(opens in a new tab)
150004
Z.ai
GLM-5.1
86.2%
Source
for GLM-5.1 on Graduate-Level Google-Proof Q&A Diamond
(opens in a new tab)
150003
Z.ai
GLM-5
(thinking)
86.0%
Source
for GLM-5 (thinking) on Graduate-Level Google-Proof Q&A Diamond
(opens in a new tab)
150002
Z.ai
GLM-4.7
85.7%
Source
for GLM-4.7 on Graduate-Level Google-Proof Q&A Diamond
(opens in a new tab)
150001
Previous
Page
1
of
1
Next