Claude Opus 4.6 · Claude Opus 4.7

Shared benchmarks

7
Shared benchmarks for Claude Opus 4.6 and Claude Opus 4.7
BenchmarkClaude Opus 4.6Claude Opus 4.7
91.3%
94.2%
Evaluation details for GPQA Diamond

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

40.0%
46.9%
Evaluation details for Humanity's Last Exam Full set — No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

53.0%
43.0%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
53.2%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
55.4%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
48.4%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
54.7%Full set — With tools · Humanity's Last Exam Accuracy
Anthropic
Evaluation details for Humanity's Last Exam Full set — With tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

76.8%
79.1%
Evaluation details for MCP Atlas Public, April 2026 update

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

77.8%
80.5%
Evaluation details for SWE-bench Multilingual

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

80.8%
87.6%
Evaluation details for SWE-bench Verified

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7

65.4%
69.4%
Evaluation details for Terminal-Bench 2.0

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

Claude Opus 4.6

Claude Opus 4.7