Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

SWE-bench

Benchmark
SWE-bench
Evaluated by
OpenAI, Princeton University
Release date
August 13, 2024
Version
SWE-bench Verified
Metric
Resolved Tasks

Results

3 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicDeepSeek AIGoogle (DeepMind)Meta AI (originally Facebook AI Research)Microsoft AIMiniMaxMistral AIMoonshot AINVIDIAOpenAIZ.ai
SWE-bench Verified results
ProviderModelScoreSourceRegistry No.
Mistral AIMistral Medium 3.5 (high)77.6%Source for Mistral Medium 3.5 (high) on SWE-bench Verified (opens in a new tab)90005
Mistral AIDevstral 272.2%Source for Devstral 2 on SWE-bench Verified (opens in a new tab)90002
Mistral AIDevstral Small 268.0%Source for Devstral Small 2 on SWE-bench Verified (opens in a new tab)90003
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme