Kimi K2.6 vs Kimi K3
12 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
12| Benchmark | Kimi K2.6 | Kimi K3 |
|---|---|---|
BabyVisionPotentially non-equivalent: Benchmark versions differ. | 39.8%Original · Accuracy Moonshot AI | 85.7%Original — Python enabled · Accuracy Moonshot AI |
Potentially non-equivalent · Evaluation details for BabyVisionBenchmark versions differ. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
83.2% | 91.2% | |
Evaluation details for BrowseCompThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
80.4% | 84.8% | |
Evaluation details for CharXiv Reasoning — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
86.7% | 91.3% | |
Evaluation details for CharXiv Reasoning — Python enabledThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
90.5% | 93.5% | |
Evaluation details for GPQA DiamondThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
34.7% | 43.5% | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
54.0% | 56.0% | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
87.4% | 94.3% | |
Evaluation details for MATH-Vision OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
79.4% | 81.6% | |
Evaluation details for MMMU-Pro No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
MMMU-ProPotentially non-equivalent: Benchmark versions differ. | 80.1%With Python · MMMU-Pro accuracy Moonshot AI | 83.4%With tools · MMMU-Pro accuracy Moonshot AI |
Potentially non-equivalent · Evaluation details for MMMU-ProBenchmark versions differ. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
73.1% | 84.8% | |
Evaluation details for OSWorld VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 66.7%2.0 · Terminal-Bench Accuracy Moonshot AI | 88.3%2.1 · Terminal-Bench Accuracy Moonshot AI |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. Kimi K2.6 Kimi K3 | ||