| | 52.5%4.0 · CursorBench Accuracy Cursor 57.8%4.0 · CursorBench Accuracy Cursor |
|---|
Evaluation details for CursorBench 4.0The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 46.6%
- Evaluator
- Cursor
- Metric
- CursorBench Accuracy (percent)
- Reasoning
- max
- Reported
- September 10, 2026
Claude Opus 5.5 - Score
- 52.5%
- Evaluator
- Cursor
- Metric
- CursorBench Accuracy (percent)
- Reasoning
- medium
- Reported
- September 22, 2026
- Score
- 57.8%
- Evaluator
- Cursor
- Metric
- CursorBench Accuracy (percent)
- Reasoning
- max
- Reported
- September 22, 2026
|
DeepSWE 1.1ⓘPotentially non-equivalent: Evaluator sets differ. | 57.7%1.1 · DeepSWE Pass@1 Anthropic 66.9%1.1 · DeepSWE Pass@1 Anthropic 69.7%1.1 · DeepSWE Pass@1 Anthropic 68.0%1.1 · DeepSWE Pass@1 Anthropic 68.8%1.1 · DeepSWE Pass@1 Anthropic 74.0%1.1 · DeepSWE Pass@1 DataCurve | 74.2%1.1 · DeepSWE Pass@1 Anthropic |
|---|
Potentially non-equivalent · Evaluation details for DeepSWE 1.1Evaluator sets differ. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 57.7%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- low
- Reported
- July 24, 2026
- Score
- 66.9%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- medium
- Reported
- July 24, 2026
- Score
- 69.7%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- xhigh
- Reported
- July 24, 2026
- Score
- 68.0%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- high
- Reported
- July 24, 2026
- Score
- 68.8%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- max
- Reported
- July 24, 2026
- Score
- 74.0%
- Evaluator
- DataCurve
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- max
- Reported
- September 22, 2026
Claude Opus 5.5 - Score
- 74.2%
- Evaluator
- Anthropic
- Metric
- DeepSWE Pass@1 (percent)
- Reasoning
- max
- Reported
- September 22, 2026
|
| | |
|---|
Evaluation details for Global MMLU OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 92.5%
- Evaluator
- Anthropic
- Metric
- Average accuracy (percent)
- Reasoning
- max
- Reported
- July 24, 2026
Claude Opus 5.5 - Score
- 94.3%
- Evaluator
- Anthropic
- Metric
- Average accuracy (percent)
- Reasoning
- max
- Reported
- September 22, 2026
|
| 47.8%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.0%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.4%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 54.2%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 56.3%Full set — No tools · Humanity's Last Exam Accuracy Anthropic | 59.0%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 64.4%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 59.6%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 52.9%Full set — No tools · Humanity's Last Exam Accuracy Anthropic 62.8%Full set — No tools · Humanity's Last Exam Accuracy Anthropic |
|---|
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 47.8%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- low
- Reported
- July 24, 2026
- Score
- 56.0%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- high
- Reported
- July 24, 2026
- Score
- 56.4%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- xhigh
- Reported
- July 24, 2026
- Score
- 54.2%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- medium
- Reported
- July 24, 2026
- Score
- 56.3%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- max
- Reported
- July 24, 2026
Claude Opus 5.5 - Score
- 59.0%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- medium
- Reported
- September 22, 2026
- Score
- 64.4%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- max
- Reported
- September 22, 2026
- Score
- 59.6%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- high
- Reported
- September 22, 2026
- Score
- 52.9%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- low
- Reported
- September 22, 2026
- Score
- 62.8%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- xhigh
- Reported
- September 22, 2026
|
| 64.8%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 64.7%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 56.1%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 63.2%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 61.3%Full set — With tools · Humanity's Last Exam Accuracy Anthropic | 63.0%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 63.9%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 66.4%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 67.7%Full set — With tools · Humanity's Last Exam Accuracy Anthropic 57.4%Full set — With tools · Humanity's Last Exam Accuracy Anthropic |
|---|
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 64.8%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- xhigh
- Reported
- July 24, 2026
- Score
- 64.7%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- max
- Reported
- July 24, 2026
- Score
- 56.1%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- low
- Reported
- July 24, 2026
- Score
- 63.2%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- high
- Reported
- July 24, 2026
- Score
- 61.3%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- medium
- Reported
- July 24, 2026
Claude Opus 5.5 - Score
- 63.0%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- medium
- Reported
- September 22, 2026
- Score
- 63.9%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- high
- Reported
- September 22, 2026
- Score
- 66.4%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- xhigh
- Reported
- September 22, 2026
- Score
- 67.7%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- adaptive thinking, max; with tools
- Reported
- September 22, 2026
- Score
- 57.4%
- Evaluator
- Anthropic
- Metric
- Humanity's Last Exam Accuracy (percent)
- Reasoning
- low
- Reported
- September 22, 2026
|
| | |
|---|
Evaluation details for MILU OriginalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 92.1%
- Evaluator
- Anthropic
- Metric
- Average accuracy (percent)
- Reasoning
- max
- Reported
- July 24, 2026
Claude Opus 5.5 - Score
- 93.1%
- Evaluator
- Anthropic
- Metric
- Average accuracy (percent)
- Reasoning
- max
- Reported
- September 22, 2026
|
| | |
|---|
Evaluation details for SWE-bench MultilingualThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. Claude Opus 5 - Score
- 89.5%
- Evaluator
- Anthropic
- Metric
- Resolved Tasks (percent)
- Reasoning
- max
- Reported
- July 24, 2026
Claude Opus 5.5 - Score
- 93.9%
- Evaluator
- Anthropic
- Metric
- Resolved Tasks (percent)
- Reasoning
- max
- Reported
- September 22, 2026
|