Claude Sonnet 4.6 vs Claude Sonnet 5
6 shared benchmark contexts with reported scores and primary source links.
Claude Sonnet 4.6 · Claude Sonnet 5
Shared benchmarks
6| Benchmark | Claude Sonnet 4.6 | Claude Sonnet 5 |
|---|---|---|
33.2% | 43.2% | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||
49.0% | 47.2%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 36.5%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 54.6%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 52.8%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||
72.5% | 81.2% | |
Evaluation details for OSWorld VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||
75.9% | 78.3% | |
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||
79.6% | 85.2% | |
Evaluation details for SWE-bench VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 59.1%2.0 · Terminal-Bench Accuracy Anthropic | 80.4%2.1 — mini-swe-agent · Terminal-Bench Accuracy Anthropic |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. Claude Sonnet 4.6 Claude Sonnet 5 | ||