Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

Humanity's Last Exam

Benchmark
Humanity's Last Exam
Evaluated by
Center for AI Safety, Scale AI
Release date
January 24, 2025
Version
Humanity's Last Exam Full set — No tools
Metric
Humanity's Last Exam Accuracy

Results

3 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicDeepSeek AIGoogle (DeepMind)Moonshot AINVIDIAOpenAI
Humanity's Last Exam Full set — No tools results
ProviderModelScoreSourceRegistry No.
Moonshot AIKimi K3 (max)43.5%Source for Kimi K3 (max) on Humanity's Last Exam Full set — No tools (opens in a new tab)120003
Moonshot AIKimi K2.6 (thinking)34.7%Source for Kimi K2.6 (thinking) on Humanity's Last Exam Full set — No tools (opens in a new tab)120002
Moonshot AIKimi K2.5 (thinking)30.1%Source for Kimi K2.5 (thinking) on Humanity's Last Exam Full set — No tools (opens in a new tab)120001
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme