Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

Toolathlon

Benchmark
Toolathlon
Evaluated by
Toolathlon Authors
Release date
June 30, 2026
Version
Toolathlon Verified
Metric
Tasks completed

Results

6 results

LatestHistory
All providersAnthropicMeta AI (originally Facebook AI Research)Thinking Machines LabZ.ai
Toolathlon Verified results
ProviderModelScoreSourceRegistry No.
AnthropicClaude Sonnet 5.5 (max)77.8%Source for Claude Sonnet 5.5 (max) on Toolathlon Verified (opens in a new tab)20016
Z.aiGLM-5.3-Flash (max)78.4%Source for GLM-5.3-Flash (max) on Toolathlon Verified (opens in a new tab)150006
Z.aiGLM-5.3 (max)73.0%Source for GLM-5.3 (max) on Toolathlon Verified (opens in a new tab)150005
Thinking Machines LabInkling-Small (effort=0.99)54.4%Source for Inkling-Small (effort=0.99) on Toolathlon Verified (opens in a new tab)160002
Thinking Machines LabInkling (effort=0.99)45.5%Source for Inkling (effort=0.99) on Toolathlon Verified (opens in a new tab)160001
Meta AI (originally Facebook AI Research)Muse Spark 1.1 (xhigh)75.6%Source for Muse Spark 1.1 (xhigh) on Toolathlon Verified (opens in a new tab)80002
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme