Qwen3.6-Plus · Qwen3.7-Plus

Shared benchmarks

13
Shared benchmarks for Qwen3.6-Plus and Qwen3.7-Plus
BenchmarkQwen3.6-PlusQwen3.7-Plus
74.2%
79.1%
Evaluation details for IFBench Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

88.0%
90.3%
Evaluation details for MATH-Vision Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

68.7%
71.0%
Evaluation details for MedXpertQA MM

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

88.5%
88.5%
Evaluation details for MMLU-Pro Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

89.5%
89.0%
Evaluation details for MMMLU Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

62.5%
73.3%
Evaluation details for OSWorld Verified

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

85.4%
86.9%
Evaluation details for RealWorldQA Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

41.4%
51.3%
Evaluation details for SciCode Original

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

45.7%
54.9%
Evaluation details for SkillsBench 78-task self-contained subset — OpenCode

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

73.8%
75.8%
Evaluation details for SWE-bench Multilingual

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

78.8%
77.7%
Evaluation details for SWE-bench Verified

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

SWE-bench ProPotentially non-equivalent: Benchmark versions differ.
56.6%Public — Qwen3.6 refined tasks · Public tasks resolved
Qwen Team
57.6%Public — Qwen refined tasks · Public tasks resolved
Qwen Team
Potentially non-equivalent · Evaluation details for SWE-bench Pro

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus

Terminal-BenchPotentially non-equivalent: Benchmark versions differ.
61.6%2.0 — Qwen3.6 Harbor/Terminus-2 · Terminal-Bench Accuracy
Qwen Team
70.3%2.0 — Qwen3.7 Harbor/Terminus-2, 5h · Terminal-Bench Accuracy
Qwen Team
Potentially non-equivalent · Evaluation details for Terminal-Bench

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

Qwen3.6-Plus

Qwen3.7-Plus