GPT-5.2 vs GPT-5.4
6 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
6| Benchmark | GPT-5.2 | GPT-5.4 |
|---|---|---|
77.9%BrowseComp · BrowseComp Accuracy OpenAI 65.8%BrowseComp · BrowseComp Accuracy OpenAI | 82.7% | |
Evaluation details for BrowseCompThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||
70.9%GDPval · Wins or ties OpenAI 74.1%GDPval · Wins or ties OpenAI | 83.0% | |
Evaluation details for GDPvalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||
93.2%Diamond · GPQA Diamond Accuracy OpenAI 92.4%Diamond · GPQA Diamond Accuracy OpenAI | 92.8% | |
Evaluation details for GPQA DiamondThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||
79.5% | 81.2% | |
Evaluation details for MMMU-Pro No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||
MMMU-ProPotentially non-equivalent: Benchmark versions differ. | 80.4%With Python · MMMU-Pro accuracy OpenAI | 82.1%With tools · MMMU-Pro accuracy OpenAI |
Potentially non-equivalent · Evaluation details for MMMU-ProBenchmark versions differ. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||
55.6% | 57.7% | |
Evaluation details for SWE-bench Pro PublicThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.2 GPT-5.4 | ||