Skip to main content
Benchmark Registry
Models
Benchmarks
Organizations
Search the registry
Search
SWE-bench
Benchmark
SWE-bench
Evaluated by
Princeton University
Release date
May 6, 2025
Version
SWE-bench Multilingual
Metric
Resolved Tasks
Search models
Search
Results
2 results
Rows per page
50
100
500
Apply
Latest
History
All providers
Alibaba (Tongyi)
Anthropic
Cursor
DeepSeek AI
MiniMax
Mistral AI
Moonshot AI
NVIDIA
Z.ai
SWE-bench Multilingual results
Provider
Model
Score
Source
Registry No.
Moonshot AI
Kimi K2.6
(thinking)
76.7%
Source
for Kimi K2.6 (thinking) on SWE-bench Multilingual
(opens in a new tab)
120002
Moonshot AI
Kimi K2.5
(non-thinking)
73.0%
Source
for Kimi K2.5 (non-thinking) on SWE-bench Multilingual
(opens in a new tab)
120001
Previous
Page
1
of
1
Next