Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

Humanity's Last Exam

Benchmark
Humanity's Last Exam
Evaluated by
Center for AI Safety, Scale AI
Release date
January 24, 2025
Version
Humanity's Last Exam Full set — No tools
Metric
Humanity's Last Exam Accuracy

Results

4 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicDeepSeek AIGoogle (DeepMind)Moonshot AINVIDIAOpenAI
Humanity's Last Exam Full set — No tools results
ProviderModelScoreSourceRegistry No.
Google (DeepMind)Gemini 3.5 Flash40.2%Source for Gemini 3.5 Flash on Humanity's Last Exam Full set — No tools (opens in a new tab)30006
Google (DeepMind)Gemini 3.1 Pro (High)44.4%Source for Gemini 3.1 Pro (High) on Humanity's Last Exam Full set — No tools (opens in a new tab)30005
Google (DeepMind)Gemini 3 Flash33.7%Source for Gemini 3 Flash on Humanity's Last Exam Full set — No tools (opens in a new tab)30009
Google (DeepMind)Gemini 3 Pro (high)37.5%Source for Gemini 3 Pro (high) on Humanity's Last Exam Full set — No tools (opens in a new tab)30008
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme