Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizations

SWE-bench

Benchmark
SWE-bench
Evaluated by
Princeton University
Release date
May 6, 2025
Version
SWE-bench Multilingual
Metric
Resolved Tasks

Results

2 results

LatestHistory
All providersAlibaba (Tongyi)AnthropicCursorDeepSeek AIMiniMaxMistral AIMoonshot AINVIDIAZ.ai
SWE-bench Multilingual results
ProviderModelScoreSourceRegistry No.
Moonshot AIKimi K2.6 (thinking)76.7%Source for Kimi K2.6 (thinking) on SWE-bench Multilingual (opens in a new tab)120002
Moonshot AIKimi K2.5 (non-thinking)73.0%Source for Kimi K2.5 (non-thinking) on SWE-bench Multilingual (opens in a new tab)120001
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme