Claude Opus 4.8 · Claude Opus 5

Shared benchmarks

6
Shared benchmarks for Claude Opus 4.8 and Claude Opus 5
BenchmarkClaude Opus 4.8Claude Opus 5
DeepSWE 1.1Potentially non-equivalent: Evaluator sets differ.
59.0%1.1 · DeepSWE Pass@1
DataCurve
57.7%1.1 · DeepSWE Pass@1
Anthropic
66.9%1.1 · DeepSWE Pass@1
Anthropic
69.7%1.1 · DeepSWE Pass@1
Anthropic
68.0%1.1 · DeepSWE Pass@1
Anthropic
68.8%1.1 · DeepSWE Pass@1
Anthropic
74.0%1.1 · DeepSWE Pass@1
DataCurve
Potentially non-equivalent · Evaluation details for DeepSWE 1.1

Evaluator sets differ. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5

49.8%
47.8%Full set — No tools · Humanity's Last Exam Accuracy
Anthropic
56.0%Full set — No tools · Humanity's Last Exam Accuracy
Anthropic
56.4%Full set — No tools · Humanity's Last Exam Accuracy
Anthropic
54.2%Full set — No tools · Humanity's Last Exam Accuracy
Anthropic
56.3%Full set — No tools · Humanity's Last Exam Accuracy
Anthropic
Evaluation details for Humanity's Last Exam Full set — No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5

57.9%
64.8%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
64.7%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
56.1%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
63.2%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
61.3%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
Evaluation details for Humanity's Last Exam Full set — With tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5

82.2%
85.8%
Evaluation details for MCP Atlas Public, April 2026 update

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5

84.4%
89.5%
Evaluation details for SWE-bench Multilingual

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5

88.6%
96.0%
Evaluation details for SWE-bench Verified

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.8

Claude Opus 5