Claude Opus 4.7 vs Claude Opus 4.8
8 shared benchmark contexts with reported scores and primary source links.
Claude Opus 4.7 · Claude Opus 4.8
Shared benchmarks
8| Benchmark | Claude Opus 4.7 | Claude Opus 4.8 |
|---|---|---|
94.2% | 93.6% | |
Evaluation details for GPQA DiamondThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
46.9% | 49.8% | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
43.0%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 53.2%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 55.4%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 48.4%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 54.7%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | 57.9% | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
79.1% | 82.2% | |
Evaluation details for MCP Atlas Public, April 2026 updateThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
80.5% | 84.4% | |
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
87.6% | 88.6% | |
Evaluation details for SWE-bench VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
64.3% | 69.2% | |
Evaluation details for SWE-bench Pro PublicThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||
Terminal-BenchPotentially non-equivalent: Benchmark versions differ. | 69.4%2.0 · Terminal-Bench Accuracy Anthropic | 74.6%2.1 · Terminal-Bench Accuracy Anthropic |
Potentially non-equivalent · Evaluation details for Terminal-BenchBenchmark versions differ. Evaluation methodology is not recorded by the registry. Claude Opus 4.7 Claude Opus 4.8 | ||