| MiniMax M3 | BrowseComp | 83.5% | minimax.io for MiniMax M3 on BrowseComp (opens in a new tab) |
| MiniMax M3 | CL-bench Original | 20.48% | minimax.io for MiniMax M3 on CL-bench (opens in a new tab) |
| Qwen3.7-Plus | RealWorldQA Original | 86.9% | qwen.ai for Qwen3.7-Plus on RealWorldQA (opens in a new tab) |
| Qwen3.7-Plus | MMMLU Original | 89.0% | qwen.ai for Qwen3.7-Plus on MMMLU (opens in a new tab) |
| Qwen3.7-Max | MMMLU Original | 90.3% | qwen.ai for Qwen3.7-Max on MMMLU (opens in a new tab) |
| Qwen3.7-Max | MMLU-Pro Original | 89.6% | qwen.ai for Qwen3.7-Max on MMLU-Pro (opens in a new tab) |
| Claude Opus 4.7 | CyberGym Original | 73.1% | anthropic.com for Claude Opus 4.7 on CyberGym (opens in a new tab) |
| Claude Opus 4.7 | MILU Original | 89.9% | www-cdn.anthropic.com for Claude Opus 4.7 on Multi-task Indic Language Understanding Benchmark (opens in a new tab) |
| Claude Opus 4.7 | Global MMLU Original | 89.9% | www-cdn.anthropic.com for Claude Opus 4.7 on Global MMLU (opens in a new tab) |
| Claude Opus 4.7 | Humanity's Last Exam Full set — With tools | 55.4% | www-cdn.anthropic.com for Claude Opus 4.7 on Humanity's Last Exam (opens in a new tab) |
| Claude Opus 4.7 | Humanity's Last Exam Full set — With tools | 53.2% | www-cdn.anthropic.com for Claude Opus 4.7 on Humanity's Last Exam (opens in a new tab) |
| Claude Opus 4.7 | Humanity's Last Exam Full set — With tools | 48.4% | www-cdn.anthropic.com for Claude Opus 4.7 on Humanity's Last Exam (opens in a new tab) |
| Claude Opus 4.7 | Humanity's Last Exam Full set — With tools | 43.0% | www-cdn.anthropic.com for Claude Opus 4.7 on Humanity's Last Exam (opens in a new tab) |
| Claude Opus 4.7 | CharXiv Reasoning — Python enabled | 91.0% | www-cdn.anthropic.com for Claude Opus 4.7 on CharXiv (opens in a new tab) |
| Claude Opus 4.7 | CharXiv Reasoning — No tools | 82.1% | www-cdn.anthropic.com for Claude Opus 4.7 on CharXiv (opens in a new tab) |
| Claude Opus 4.7 | MMMLU Original | 91.5% | www-cdn.anthropic.com for Claude Opus 4.7 on MMMLU (opens in a new tab) |
| Qwen3.6-Plus | MMMLU Original | 89.5% | qwen.ai for Qwen3.6-Plus on MMMLU (opens in a new tab) |
| Qwen3.6-Plus | AI2D TEST | 94.4% | qwen.ai for Qwen3.6-Plus on AI2D (opens in a new tab) |
| Qwen3.6-Plus | RealWorldQA Original | 85.4% | qwen.ai for Qwen3.6-Plus on RealWorldQA (opens in a new tab) |
| Qwen3.6-Plus | LongBench v2 | 62.0% | qwen.ai for Qwen3.6-Plus on LongBench (opens in a new tab) |
| Claude Fable 5 | OSWorld Verified | 85.0% | www-cdn.anthropic.com for Claude Fable 5 on OSWorld (opens in a new tab) |
| Claude Fable 5 | Terminal-Bench 2.1 — mini-swe-agent | 84.3% | www-cdn.anthropic.com for Claude Fable 5 on Terminal-Bench (opens in a new tab) |
| Claude Fable 5 | SWE-bench Verified | 95.0% | www-cdn.anthropic.com for Claude Fable 5 on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Pro | Toolathlon Original | 49.0% | huggingface.co for DeepSeek-V4-Pro on Toolathlon (opens in a new tab) |
| DeepSeek-V4-Pro | Toolathlon Original | 46.3% | huggingface.co for DeepSeek-V4-Pro on Toolathlon (opens in a new tab) |
| DeepSeek-V4-Flash | Toolathlon Original | 43.5% | huggingface.co for DeepSeek-V4-Flash on Toolathlon (opens in a new tab) |
| DeepSeek-V4-Flash | Toolathlon Original | 40.7% | huggingface.co for DeepSeek-V4-Flash on Toolathlon (opens in a new tab) |
| DeepSeek-V4-Pro | BrowseComp | 80.4% | huggingface.co for DeepSeek-V4-Pro on BrowseComp (opens in a new tab) |
| DeepSeek-V4-Flash | BrowseComp | 53.5% | huggingface.co for DeepSeek-V4-Flash on BrowseComp (opens in a new tab) |
| DeepSeek-V4-Pro | SWE-bench Multilingual | 74.1% | huggingface.co for DeepSeek-V4-Pro on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Pro | SWE-bench Multilingual | 69.8% | huggingface.co for DeepSeek-V4-Pro on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Flash | SWE-bench Multilingual | 70.2% | huggingface.co for DeepSeek-V4-Flash on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Flash | SWE-bench Multilingual | 69.7% | huggingface.co for DeepSeek-V4-Flash on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Pro | SWE-bench Verified | 79.4% | huggingface.co for DeepSeek-V4-Pro on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Pro | SWE-bench Verified | 73.6% | huggingface.co for DeepSeek-V4-Pro on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Flash | SWE-bench Verified | 78.6% | huggingface.co for DeepSeek-V4-Flash on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Flash | SWE-bench Verified | 73.7% | huggingface.co for DeepSeek-V4-Flash on SWE-bench (opens in a new tab) |
| DeepSeek-V4-Pro | Terminal-Bench 2.0 | 63.3% | huggingface.co for DeepSeek-V4-Pro on Terminal-Bench (opens in a new tab) |
| DeepSeek-V4-Pro | Terminal-Bench 2.0 | 59.1% | huggingface.co for DeepSeek-V4-Pro on Terminal-Bench (opens in a new tab) |
| DeepSeek-V4-Flash | Terminal-Bench 2.0 | 56.6% | huggingface.co for DeepSeek-V4-Flash on Terminal-Bench (opens in a new tab) |
| DeepSeek-V4-Flash | Terminal-Bench 2.0 | 49.1% | huggingface.co for DeepSeek-V4-Flash on Terminal-Bench (opens in a new tab) |
| DeepSeek-V4-Pro | GPQA Diamond | 89.1% | huggingface.co for DeepSeek-V4-Pro on Graduate-Level Google-Proof Q&A (opens in a new tab) |
| DeepSeek-V4-Pro | GPQA Diamond | 72.9% | huggingface.co for DeepSeek-V4-Pro on Graduate-Level Google-Proof Q&A (opens in a new tab) |
| DeepSeek-V4-Flash | GPQA Diamond | 87.4% | huggingface.co for DeepSeek-V4-Flash on Graduate-Level Google-Proof Q&A (opens in a new tab) |
| DeepSeek-V4-Flash | GPQA Diamond | 71.2% | huggingface.co for DeepSeek-V4-Flash on Graduate-Level Google-Proof Q&A (opens in a new tab) |
| DeepSeek-V4-Pro | SimpleQA Verified Original | 57.9% | huggingface.co for DeepSeek-V4-Pro on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Pro | SimpleQA Verified Original | 46.2% | huggingface.co for DeepSeek-V4-Pro on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Pro | SimpleQA Verified Original | 45.0% | huggingface.co for DeepSeek-V4-Pro on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Flash | SimpleQA Verified Original | 34.1% | huggingface.co for DeepSeek-V4-Flash on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Flash | SimpleQA Verified Original | 28.9% | huggingface.co for DeepSeek-V4-Flash on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Flash | SimpleQA Verified Original | 23.1% | huggingface.co for DeepSeek-V4-Flash on SimpleQA Verified (opens in a new tab) |
| DeepSeek-V4-Pro | MMLU-Pro Original | 87.5% | huggingface.co for DeepSeek-V4-Pro on MMLU-Pro (opens in a new tab) |
| DeepSeek-V4-Pro | MMLU-Pro Original | 87.1% | huggingface.co for DeepSeek-V4-Pro on MMLU-Pro (opens in a new tab) |
| DeepSeek-V4-Pro | MMLU-Pro Original | 82.9% | huggingface.co for DeepSeek-V4-Pro on MMLU-Pro (opens in a new tab) |
| DeepSeek-V4-Flash | MMLU-Pro Original | 86.2% | huggingface.co for DeepSeek-V4-Flash on MMLU-Pro (opens in a new tab) |
| DeepSeek-V4-Flash | MMLU-Pro Original | 86.4% | huggingface.co for DeepSeek-V4-Flash on MMLU-Pro (opens in a new tab) |
| DeepSeek-V4-Flash | MMLU-Pro Original | 83.0% | huggingface.co for DeepSeek-V4-Flash on MMLU-Pro (opens in a new tab) |
| Composer 2.5 | SWE-bench Multilingual — Cursor strict harness | 71.6% | cursor.com for Composer 2.5 on SWE-bench (opens in a new tab) |
| Composer 2.5 | SWE-bench Multilingual — Cursor standard harness | 79.2% | cursor.com for Composer 2.5 on SWE-bench (opens in a new tab) |
| Muse Spark | CyberGym Original | 43.5% | ai.meta.com for Muse Spark on CyberGym (opens in a new tab) |
| Muse Spark | FrontierScience Research | 38.3% | ai.meta.com for Muse Spark on FrontierScience (opens in a new tab) |
| Muse Spark | International Physics Olympiad 2025 Theory | 82.6% | ai.meta.com for Muse Spark on International Physics Olympiad (opens in a new tab) |
| Muse Spark | Humanity's Last Exam Full set — With tools | 58.4% | ai.meta.com for Muse Spark on Humanity's Last Exam (opens in a new tab) |
| Muse Spark | Humanity's Last Exam Full set — No tools | 50.2% | ai.meta.com for Muse Spark on Humanity's Last Exam (opens in a new tab) |
| Muse Spark | Terminal-Bench 2.0 | 59.0% | ai.meta.com for Muse Spark on Terminal-Bench (opens in a new tab) |
| Muse Spark | SWE-bench Pro Public | 52.4% | ai.meta.com for Muse Spark on SWE-bench Pro (opens in a new tab) |
| Muse Spark | SWE-bench Verified — Muse Spark corrected tests | 77.4% | ai.meta.com for Muse Spark on SWE-bench (opens in a new tab) |