Qwen3.6-Plus vs Qwen3.7-Plus
13 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
13| Benchmark | Qwen3.6-Plus | Qwen3.7-Plus |
|---|---|---|
74.2% | 79.1% | |
Evaluation details for IFBench OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
88.0% | 90.3% | |
Evaluation details for MATH-Vision OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
68.7% | 71.0% | |
Evaluation details for MedXpertQA MMThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
88.5% | 88.5% | |
Evaluation details for MMLU-Pro OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
89.5% | 89.0% | |
Evaluation details for MMMLU OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
62.5% | 73.3% | |
Evaluation details for OSWorld VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
85.4% | 86.9% | |
Evaluation details for RealWorldQA OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
41.4% | 51.3% | |
Evaluation details for SciCode OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
45.7% | 54.9% | |
Evaluation details for SkillsBench 78-task self-contained subset — OpenCodeThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
73.8% | 75.8% | |
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
78.8% | 77.7% | |
Evaluation details for SWE-bench VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
SWE-bench ProPotentially non-equivalent: Benchmark versions differ. | 56.6%Public — Qwen3.6 refined tasks · Public tasks resolved Qwen Team | 57.6%Public — Qwen refined tasks · Public tasks resolved Qwen Team |
Potentially non-equivalent · Evaluation details for SWE-bench ProBenchmark versions differ. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 61.6%2.0 — Qwen3.6 Harbor/Terminus-2 · Terminal-Bench Accuracy Qwen Team | 70.3%2.0 — Qwen3.7 Harbor/Terminus-2, 5h · Terminal-Bench Accuracy Qwen Team |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. Qwen3.6-Plus Qwen3.7-Plus | ||