GPT-5.2 · GPT-5.4

Shared benchmarks

6
Shared benchmarks for GPT-5.2 and GPT-5.4
BenchmarkGPT-5.2GPT-5.4
77.9%BrowseComp · BrowseComp Accuracy
OpenAI
65.8%BrowseComp · BrowseComp Accuracy
OpenAI
82.7%
Evaluation details for BrowseComp

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4

70.9%GDPval · Wins or ties
OpenAI
74.1%GDPval · Wins or ties
OpenAI
83.0%
Evaluation details for GDPval

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4

93.2%Diamond · GPQA Diamond Accuracy
OpenAI
92.4%Diamond · GPQA Diamond Accuracy
OpenAI
92.8%
Evaluation details for GPQA Diamond

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4

79.5%
81.2%
Evaluation details for MMMU-Pro No tools

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4

MMMU-ProPotentially non-equivalent: Benchmark versions differ.
80.4%With Python · MMMU-Pro accuracy
OpenAI
82.1%With tools · MMMU-Pro accuracy
OpenAI
Potentially non-equivalent · Evaluation details for MMMU-Pro

Benchmark versions differ. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4

55.6%
57.7%
Evaluation details for SWE-bench Pro Public

The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry.

GPT-5.2

GPT-5.4