Claude Opus 4.8 vs Claude Opus 5
6 shared benchmark contexts with reported scores and primary source links.
Claude Opus 4.8 · Claude Opus 5
Shared benchmarks
6| Benchmark | Claude Opus 4.8 | Claude Opus 5 |
|---|---|---|
DeepSWE 1.1Potentially non-equivalent: Evaluator sets differ. | 59.0%1.1 · DeepSWE Pass@1 DataCurve | 57.7%1.1 · DeepSWE Pass@1 Anthropic 66.9%1.1 · DeepSWE Pass@1 Anthropic 69.7%1.1 · DeepSWE Pass@1 Anthropic 68.0%1.1 · DeepSWE Pass@1 Anthropic 68.8%1.1 · DeepSWE Pass@1 Anthropic 74.0%1.1 · DeepSWE Pass@1 DataCurve |
Potentially non-equivalent · Evaluation details for DeepSWE 1.1Evaluator sets differ. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||
49.8% | 47.8%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.0%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.4%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 54.2%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.3%Full set — No tools · Humanity's Last Exam Accuracy Anthropic | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||
57.9% | 64.8%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 64.7%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 56.1%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 63.2%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 61.3%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||
82.2% | 85.8% | |
Evaluation details for MCP Atlas Public, April 2026 updateThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||
84.4% | 89.5% | |
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||
88.6% | 96.0% | |
Evaluation details for SWE-bench VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.8 Claude Opus 5 | ||