| Microsoft AI | MAI-Thinking-1 | 73.5% | Source for MAI-Thinking-1 on SWE-bench Verified (opens in a new tab) | 70003 |
| Meta AI (originally Facebook AI Research) | Muse Glimmer 30B (high) | 76.0% | Source for Muse Glimmer 30B (high) on SWE-bench Verified (opens in a new tab) | 80004 |
| NVIDIA | NVIDIA Nemotron 3 Ultra 550B-A55B | 70.7% | Source for NVIDIA Nemotron 3 Ultra 550B-A55B on SWE-bench Verified (opens in a new tab) | 60003 |
| Alibaba (Tongyi) | Qwen3.7-Plus | 77.7% | Source for Qwen3.7-Plus on SWE-bench Verified (opens in a new tab) | 130007 |
| MiniMax | MiniMax M3 | 80.5% | Source for MiniMax M3 on SWE-bench Verified (opens in a new tab) | 140003 |
| Anthropic | Claude Opus 4.8 (adaptive thinking, max) | 88.6% | Source for Claude Opus 4.8 (adaptive thinking, max) on SWE-bench Verified (opens in a new tab) | 20011 |
| Mistral AI | Mistral Medium 3.5 (high) | 77.6% | Source for Mistral Medium 3.5 (high) on SWE-bench Verified (opens in a new tab) | 90005 |
| Alibaba (Tongyi) | Qwen3.7-Max | 80.4% | Source for Qwen3.7-Max on SWE-bench Verified (opens in a new tab) | 130006 |
| DeepSeek AI | DeepSeek-V4-Flash (max) | 79.0% | Source for DeepSeek-V4-Flash (max) on SWE-bench Verified (opens in a new tab) | 110001 |
| DeepSeek AI | DeepSeek-V4-Pro (max) | 80.6% | Source for DeepSeek-V4-Pro (max) on SWE-bench Verified (opens in a new tab) | 110002 |
| Moonshot AI | Kimi K2.6 (thinking) | 80.2% | Source for Kimi K2.6 (thinking) on SWE-bench Verified (opens in a new tab) | 120002 |
| Alibaba (Tongyi) | Qwen3.6-35B-A3B | 73.4% | Source for Qwen3.6-35B-A3B on SWE-bench Verified (opens in a new tab) | 130005 |
| Anthropic | Claude Opus 4.7 (adaptive thinking, max) | 87.6% | Source for Claude Opus 4.7 (adaptive thinking, max) on SWE-bench Verified (opens in a new tab) | 20010 |
| Alibaba (Tongyi) | Qwen3.6-Plus | 78.8% | Source for Qwen3.6-Plus on SWE-bench Verified (opens in a new tab) | 130004 |
| NVIDIA | NVIDIA Nemotron 3 Super 120B-A12B (OpenHands) | 60.5% | Source for NVIDIA Nemotron 3 Super 120B-A12B (OpenHands) on SWE-bench Verified (opens in a new tab) | 60002 |
| Google (DeepMind) | Gemini 3.1 Pro (High) | 80.6% | Source for Gemini 3.1 Pro (High) on SWE-bench Verified (opens in a new tab) | 30005 |
| Alibaba (Tongyi) | Qwen3.5-397B-A17B | 76.4% | Source for Qwen3.5-397B-A17B on SWE-bench Verified (opens in a new tab) | 130003 |
| Anthropic | Claude Sonnet 4.6 (adaptive thinking, max) | 79.6% | Source for Claude Sonnet 4.6 (adaptive thinking, max) on SWE-bench Verified (opens in a new tab) | 20009 |
| MiniMax | MiniMax M2.5 | 80.2% | Source for MiniMax M2.5 on SWE-bench Verified (opens in a new tab) | 140001 |
| Z.ai | GLM-5 (thinking) | 77.8% | Source for GLM-5 (thinking) on SWE-bench Verified (opens in a new tab) | 150002 |
| Anthropic | Claude Opus 4.6 (adaptive thinking, max) | 80.8% | Source for Claude Opus 4.6 (adaptive thinking, max) on SWE-bench Verified (opens in a new tab) | 20008 |
| Moonshot AI | Kimi K2.5 (non-thinking) | 76.8% | Source for Kimi K2.5 (non-thinking) on SWE-bench Verified (opens in a new tab) | 120001 |
| Alibaba (Tongyi) | Qwen3-Max-Thinking | 75.3% | Source for Qwen3-Max-Thinking on SWE-bench Verified (opens in a new tab) | 130002 |
| Z.ai | GLM-4.7 | 73.8% | Source for GLM-4.7 on SWE-bench Verified (opens in a new tab) | 150001 |
| Google (DeepMind) | Gemini 3 Flash | 78.0% | Source for Gemini 3 Flash on SWE-bench Verified (opens in a new tab) | 30009 |
| OpenAI | GPT-5.2 (xhigh) | 80.0% | Source for GPT-5.2 (xhigh) on SWE-bench Verified (opens in a new tab) | 10012 |
| Mistral AI | Devstral 2 | 72.2% | Source for Devstral 2 on SWE-bench Verified (opens in a new tab) | 90002 |
| Mistral AI | Devstral Small 2 | 68.0% | Source for Devstral Small 2 on SWE-bench Verified (opens in a new tab) | 90003 |
| Anthropic | Claude Opus 4.5 (no thinking) | 80.9% | Source for Claude Opus 4.5 (no thinking) on SWE-bench Verified (opens in a new tab) | 20007 |
| Google (DeepMind) | Gemini 3 Pro (high) | 76.2% | Source for Gemini 3 Pro (high) on SWE-bench Verified (opens in a new tab) | 30008 |
| Anthropic | Claude Haiku 4.5 (extended thinking (128K)) | 73.3% | Source for Claude Haiku 4.5 (extended thinking (128K)) on SWE-bench Verified (opens in a new tab) | 20006 |
| Anthropic | Claude Sonnet 4.5 (extended thinking (200K)) | 77.2% | Source for Claude Sonnet 4.5 (extended thinking (200K)) on SWE-bench Verified (opens in a new tab) | 20005 |
| Alibaba (Tongyi) | Qwen3-Max-Instruct | 69.6% | Source for Qwen3-Max-Instruct on SWE-bench Verified (opens in a new tab) | 130001 |
| OpenAI | gpt-oss-120b (high) | 62.4% | Source for gpt-oss-120b (high) on SWE-bench Verified (opens in a new tab) | 15001 |
| OpenAI | gpt-oss-120b (medium) | 52.6% | Source for gpt-oss-120b (medium) on SWE-bench Verified (opens in a new tab) | 15001 |
| OpenAI | gpt-oss-120b (low) | 47.9% | Source for gpt-oss-120b (low) on SWE-bench Verified (opens in a new tab) | 15001 |
| Google (DeepMind) | Gemini 2.0 Flash | 21.4% | Source for Gemini 2.0 Flash on SWE-bench Verified (opens in a new tab) | 30001 |
| Google (DeepMind) | Gemini 2.5 Pro | 59.6% | Source for Gemini 2.5 Pro on SWE-bench Verified (opens in a new tab) | 30002 |
| Google (DeepMind) | Gemini 2.5 Flash | 48.9% | Source for Gemini 2.5 Flash on SWE-bench Verified (opens in a new tab) | 30003 |
| Anthropic | Claude Opus 4 (standard (no extended thinking)) | 72.5% | Source for Claude Opus 4 (standard (no extended thinking)) on SWE-bench Verified (opens in a new tab) | 20002 |
| Anthropic | Claude Sonnet 4 (standard (no extended thinking)) | 72.7% | Source for Claude Sonnet 4 (standard (no extended thinking)) on SWE-bench Verified (opens in a new tab) | 20003 |
| OpenAI | GPT-4.5 Preview | 38.0% | Source for GPT-4.5 Preview on SWE-bench Verified (opens in a new tab) | 10001 |
| OpenAI | GPT-4.1 | 54.6% | Source for GPT-4.1 on SWE-bench Verified (opens in a new tab) | 10002 |
| Anthropic | Claude 3.7 Sonnet (No extended thinking) | 62.3% | Source for Claude 3.7 Sonnet (No extended thinking) on SWE-bench Verified (opens in a new tab) | 20001 |