Claude Sonnet 5 vs Claude Sonnet 5.5
8 shared benchmark contexts with reported scores and primary source links.
Claude Sonnet 5 · Claude Sonnet 5.5
Shared benchmarks
8| Benchmark | Claude Sonnet 5 | Claude Sonnet 5.5 |
|---|---|---|
34.1% | 53.1%4.0 · CursorBench Accuracy Cursor 35.8%4.0 · CursorBench Accuracy Cursor 55.5%4.0 · CursorBench Accuracy Cursor 39.2%4.0 · CursorBench Accuracy Cursor 47.8%4.0 · CursorBench Accuracy Cursor | |
Evaluation details for CursorBench 4.0The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
DeepSWE 1.1Potentially non-equivalent: Evaluator sets differ. | 54.0%1.1 · DeepSWE Pass@1 DataCurve | 71.0%1.1 · DeepSWE Pass@1 Anthropic |
Potentially non-equivalent · Evaluation details for DeepSWE 1.1Evaluator sets differ. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
89.0% | 92.1% | |
Evaluation details for Global MMLU OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
43.2% | 43.2%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.9%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 41.4%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 47.9%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 53.0%Full set — No tools · Humanity's Last Exam Accuracy Anthropic | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
47.2%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 36.5%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 54.6%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 52.8%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | 64.5%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 56.7%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 50.5%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 62.0%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 47.6%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
89.3% | 91.6% | |
Evaluation details for MILU OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
78.3% | 90.3% | |
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||
63.2% | 81.3% | |
Evaluation details for SWE-bench Pro PublicThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Sonnet 5 Claude Sonnet 5.5 | ||