Skip to main content
Benchmark Registry
Models
Benchmarks
Organizations
Search the registry
Search
SWE-bench
Benchmark
SWE-bench
Evaluated by
OpenAI, Princeton University
Release date
August 13, 2024
Version
SWE-bench Verified
Metric
Resolved Tasks
Search models
Search
Results
3 results
Rows per page
50
100
500
Apply
Latest
History
All providers
Alibaba (Tongyi)
Anthropic
DeepSeek AI
Google (DeepMind)
Meta AI (originally Facebook AI Research)
Microsoft AI
MiniMax
Mistral AI
Moonshot AI
NVIDIA
OpenAI
Z.ai
SWE-bench Verified results
Provider
Model
Score
Source
Registry No.
Mistral AI
Mistral Medium 3.5
(high)
77.6%
Source
for Mistral Medium 3.5 (high) on SWE-bench Verified
(opens in a new tab)
90005
Mistral AI
Devstral 2
72.2%
Source
for Devstral 2 on SWE-bench Verified
(opens in a new tab)
90002
Mistral AI
Devstral Small 2
68.0%
Source
for Devstral Small 2 on SWE-bench Verified
(opens in a new tab)
90003
Previous
Page
1
of
1
Next