Skip to main content
Benchmark Registry
Models
Benchmarks
Organizations
Search the registry
Search
Humanity's Last Exam
Benchmark
Humanity's Last Exam
Evaluated by
Center for AI Safety, Scale AI
Release date
January 24, 2025
Version
Humanity's Last Exam Full set — No tools
Metric
Humanity's Last Exam Accuracy
Search models
Search
Results
3 results
Rows per page
50
100
500
Apply
Latest
History
All providers
Alibaba (Tongyi)
Anthropic
DeepSeek AI
Google (DeepMind)
Moonshot AI
NVIDIA
OpenAI
Humanity's Last Exam Full set — No tools results
Provider
Model
Score
Source
Registry No.
Alibaba (Tongyi)
Qwen3.8-Max
43.6%
Source
for Qwen3.8-Max on Humanity's Last Exam Full set — No tools
(opens in a new tab)
130008
Alibaba (Tongyi)
Qwen3.5-397B-A17B
28.7%
Source
for Qwen3.5-397B-A17B on Humanity's Last Exam Full set — No tools
(opens in a new tab)
130003
Alibaba (Tongyi)
Qwen3-Max-Thinking
30.2%
Source
for Qwen3-Max-Thinking on Humanity's Last Exam Full set — No tools
(opens in a new tab)
130002
Previous
Page
1
of
1
Next