Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

SWE-bench

Benchmark
SWE-bench
Evaluated by
Princeton University
Release date
May 6, 2025
Version
SWE-bench Multilingual
Metric
Resolved Tasks

Results

5 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicCursorDeepSeek AIMiniMaxMistral AIMoonshot AINVIDIAZ.ai
SWE-bench Multilingual results
ProviderModelScoreSourceRegistry No.
AnthropicClaude Sonnet 5.5 (max)90.3%Source for Claude Sonnet 5.5 (max) on SWE-bench Multilingual (opens in a new tab)20016
AnthropicClaude Opus 4.8 (adaptive thinking, max)84.4%Source for Claude Opus 4.8 (adaptive thinking, max) on SWE-bench Multilingual (opens in a new tab)20011
AnthropicClaude Opus 4.7 (adaptive thinking, max)80.5%Source for Claude Opus 4.7 (adaptive thinking, max) on SWE-bench Multilingual (opens in a new tab)20010
AnthropicClaude Sonnet 4.6 (adaptive thinking, max)75.9%Source for Claude Sonnet 4.6 (adaptive thinking, max) on SWE-bench Multilingual (opens in a new tab)20009
AnthropicClaude Opus 4.6 (adaptive thinking, max)77.8%Source for Claude Opus 4.6 (adaptive thinking, max) on SWE-bench Multilingual (opens in a new tab)20008
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme