Skip to main content
Benchmark Registry
Models
Benchmarks
Organizations
Compare
Search the registry
Search
MMLU-Pro
Benchmark
MMLU-Pro
Evaluated by
Qwen Team
Release date
June 3, 2024
Version
MMLU-Pro Original
Metric
Accuracy
Search models
Search
Results
6 results
Rows per page
50
100
500
Apply
Latest
History
All providers
Alibaba (Tongyi)
DeepSeek AI
Google (DeepMind)
Microsoft AI
Mistral AI
Moonshot AI
NVIDIA
Z.ai
MMLU-Pro Original results
Provider
Model
Score
Source
Registry No.
Alibaba (Tongyi)
Qwen3.7-Plus
88.5%
Source
for Qwen3.7-Plus on MMLU-Pro Original
(opens in a new tab)
130007
Alibaba (Tongyi)
Qwen3.7-Max
89.6%
Source
for Qwen3.7-Max on MMLU-Pro Original
(opens in a new tab)
130006
Alibaba (Tongyi)
Qwen3.6-35B-A3B
85.2%
Source
for Qwen3.6-35B-A3B on MMLU-Pro Original
(opens in a new tab)
130005
Alibaba (Tongyi)
Qwen3.6-Plus
88.5%
Source
for Qwen3.6-Plus on MMLU-Pro Original
(opens in a new tab)
130004
Alibaba (Tongyi)
Qwen3.5-397B-A17B
87.8%
Source
for Qwen3.5-397B-A17B on MMLU-Pro Original
(opens in a new tab)
130003
Alibaba (Tongyi)
Qwen3-Max-Thinking
85.7%
Source
for Qwen3-Max-Thinking on MMLU-Pro Original
(opens in a new tab)
130002
Previous
Page
1
of
1
Next