Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizationsCompare

MMLU-Pro

Benchmark
MMLU-Pro
Evaluated by
Qwen Team
Release date
June 3, 2024
Version
MMLU-Pro Original
Metric
Accuracy

Results

5 results

LatestHistory
All providersAlibaba (Tongyi)DeepSeek AIGoogle (DeepMind)Microsoft AIMistral AIMoonshot AINVIDIAZ.ai
MMLU-Pro Original results
ProviderModelScoreSourceRegistry No.
Google (DeepMind)Gemma 4 26B A4B Instruct (thinking mode)82.6%Source for Gemma 4 26B A4B Instruct (thinking mode) on MMLU-Pro Original (opens in a new tab)35001
Google (DeepMind)Gemma 4 31B Instruct (thinking mode)85.2%Source for Gemma 4 31B Instruct (thinking mode) on MMLU-Pro Original (opens in a new tab)35002
Google (DeepMind)Gemma 4 E2B Instruct (thinking mode)60.0%Source for Gemma 4 E2B Instruct (thinking mode) on MMLU-Pro Original (opens in a new tab)35003
Google (DeepMind)Gemma 4 E4B Instruct (thinking mode)69.4%Source for Gemma 4 E4B Instruct (thinking mode) on MMLU-Pro Original (opens in a new tab)35004
Google (DeepMind)Gemma 4 12B Instruct (thinking mode)77.2%Source for Gemma 4 12B Instruct (thinking mode) on MMLU-Pro Original (opens in a new tab)35005
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme