Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

SWE-bench Pro

Benchmark
SWE-bench Pro
Evaluated by
Scale AI
Release date
September 19, 2025
Version
SWE-bench Pro Public
Metric
Public tasks resolved

Results

3 results

LatestHistory
All providersAnthropicDeepSeek AIGoogle (DeepMind)Meta AI (originally Facebook AI Research)Microsoft AIMiniMaxMoonshot AIOpenAISpaceXAIThinking Machines LabZ.ai
SWE-bench Pro Public results
ProviderModelScoreSourceRegistry No.
AnthropicClaude Sonnet 5.5 (max)81.3%Source for Claude Sonnet 5.5 (max) on SWE-bench Pro Public (opens in a new tab)20016
AnthropicClaude Opus 4.8 (adaptive thinking, max)69.2%Source for Claude Opus 4.8 (adaptive thinking, max) on SWE-bench Pro Public (opens in a new tab)20011
AnthropicClaude Opus 4.7 (adaptive thinking, max)64.3%Source for Claude Opus 4.7 (adaptive thinking, max) on SWE-bench Pro Public (opens in a new tab)20010
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme