Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

OSWorld

Benchmark
OSWorld
Evaluated by
XLang Lab
Release date
July 28, 2025
Version
OSWorld Verified
Metric
OSWorld-Verified success rate

Results

5 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicGoogle (DeepMind)Meta AI (originally Facebook AI Research)Moonshot AIOpenAI
OSWorld Verified results
ProviderModelScoreSourceRegistry No.
AnthropicClaude Sonnet 5 (max)81.2%Source for Claude Sonnet 5 (max) on OSWorld Verified (opens in a new tab)20013
AnthropicClaude Opus 4.8 (adaptive thinking, max)83.4%Source for Claude Opus 4.8 (adaptive thinking, max) on OSWorld Verified (opens in a new tab)20011
AnthropicClaude Sonnet 4.6 (adaptive thinking, max)72.5%Source for Claude Sonnet 4.6 (adaptive thinking, max) on OSWorld Verified (opens in a new tab)20009
AnthropicClaude Opus 4.6 (adaptive thinking, max)72.7%Source for Claude Opus 4.6 (adaptive thinking, max) on OSWorld Verified (opens in a new tab)20008
AnthropicClaude Opus 4.5 (extended thinking)66.3%Source for Claude Opus 4.5 (extended thinking) on OSWorld Verified (opens in a new tab)20007
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme