Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

Humanity's Last Exam

Benchmark
Humanity's Last Exam
Evaluated by
Center for AI Safety, Scale AI
Release date
January 24, 2025
Version
Humanity's Last Exam Full set — With tools
Metric
Humanity's Last Exam Accuracy

Results

3 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicDeepSeek AIMeta AI (originally Facebook AI Research)Moonshot AINVIDIAOpenAIZ.ai
Humanity's Last Exam Full set — With tools results
ProviderModelScoreSourceRegistry No.
Alibaba (Tongyi)Qwen3.8-Max56.2%Source for Qwen3.8-Max on Humanity's Last Exam Full set — With tools (opens in a new tab)130008
Alibaba (Tongyi)Qwen3.5-397B-A17B48.3%Source for Qwen3.5-397B-A17B on Humanity's Last Exam Full set — With tools (opens in a new tab)130003
Alibaba (Tongyi)Qwen3-Max-Thinking49.8%Source for Qwen3-Max-Thinking on Humanity's Last Exam Full set — With tools (opens in a new tab)130002
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme