Kimi K2.6 · Kimi K3

Shared benchmarks

12
Shared benchmarks for Kimi K2.6 and Kimi K3
BenchmarkKimi K2.6Kimi K3
BabyVisionPotentially non-equivalent: Benchmark versions differ.
39.8%Original · Accuracy
Moonshot AI
85.7%Original — Python enabled · Accuracy
Moonshot AI
Potentially non-equivalent · Evaluation details for BabyVision

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

83.2%
91.2%
Evaluation details for BrowseComp

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

80.4%
84.8%
Evaluation details for CharXiv Reasoning — No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

86.7%
91.3%
Evaluation details for CharXiv Reasoning — Python enabled

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

90.5%
93.5%
Evaluation details for GPQA Diamond

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

34.7%
43.5%
Evaluation details for Humanity's Last Exam Full set — No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

54.0%
56.0%
Evaluation details for Humanity's Last Exam Full set — With tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

87.4%
94.3%
Evaluation details for MATH-Vision Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

79.4%
81.6%
Evaluation details for MMMU-Pro No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

MMMU-ProPotentially non-equivalent: Benchmark versions differ.
80.1%With Python · MMMU-Pro accuracy
Moonshot AI
83.4%With tools · MMMU-Pro accuracy
Moonshot AI
Potentially non-equivalent · Evaluation details for MMMU-Pro

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

73.1%
84.8%
Evaluation details for OSWorld Verified

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3

Terminal-BenchPotentially non-equivalent: Benchmark versions differ.
66.7%2.0 · Terminal-Bench Accuracy
Moonshot AI
88.3%2.1 · Terminal-Bench Accuracy
Moonshot AI
Potentially non-equivalent · Evaluation details for Terminal-Bench

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Kimi K2.6

Kimi K3