Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

Humanity's Last Exam

Benchmark
Humanity's Last Exam
Evaluated by
Center for AI Safety, Scale AI
Release date
January 24, 2025
Version
Humanity's Last Exam Full set — With tools
Metric
Humanity's Last Exam Accuracy

Results

2 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicDeepSeek AIMeta AI (originally Facebook AI Research)Moonshot AINVIDIAOpenAIZ.ai
Humanity's Last Exam Full set — With tools results
ProviderModelScoreSourceRegistry No.
OpenAIGPT-5.5 (xhigh)52.2%Source for GPT-5.5 (xhigh) on Humanity's Last Exam Full set — With tools (opens in a new tab)10008
OpenAIGPT-5.4 (xhigh)52.1%Source for GPT-5.4 (xhigh) on Humanity's Last Exam Full set — With tools (opens in a new tab)10007
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme