GPT-5.4 vs GPT-5.5
10 shared benchmark contexts with reported scores and primary source links.
Shared benchmarks
10| Benchmark | GPT-5.4 | GPT-5.5 |
|---|---|---|
82.7% | 84.4% | |
Evaluation details for BrowseCompThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
83.0% | 84.9% | |
Evaluation details for GDPvalThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
92.8% | 93.6% | |
Evaluation details for GPQA DiamondThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
39.8% | 41.4% | |
Evaluation details for Humanity's Last Exam Full set — No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
52.1% | 52.2% | |
Evaluation details for Humanity's Last Exam Full set — With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
70.6% | 75.3% | |
Evaluation details for MCP Atlas Public, April 2026 updateThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
81.2% | 81.2% | |
Evaluation details for MMMU-Pro No toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
82.1% | 83.2% | |
Evaluation details for MMMU-Pro With toolsThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
75.0% | 78.7% | |
Evaluation details for OSWorld VerifiedThe recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||
75.1% | 82.7% | |
Evaluation details for Terminal-Bench 2.0The recorded benchmark version, metric, and evaluator sets match. Evaluation methodology is not recorded by the registry. GPT-5.4 GPT-5.5 | ||