Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizationsCompare

Terminal-Bench

Benchmark
Terminal-Bench
Evaluated by
Terminal-Bench Team
Release date
November 7, 2025
Version
Terminal-Bench 2.0
Metric
Terminal-Bench Accuracy

Results

3 results

LatestHistory
All providersAnthropicCursorDeepSeek AIGoogle (DeepMind)Meta AI (originally Facebook AI Research)Microsoft AIMiniMaxMistral AIMoonshot AINVIDIAOpenAIZ.ai
Terminal-Bench 2.0 results
ProviderModelScoreSourceRegistry No.
OpenAIGPT-5.5 (xhigh)82.7%Source for GPT-5.5 (xhigh) on Terminal-Bench 2.0 (opens in a new tab)10008
OpenAIGPT-5.4 (xhigh)75.1%Source for GPT-5.4 (xhigh) on Terminal-Bench 2.0 (opens in a new tab)10007
OpenAIGPT-5.3-Codex (xhigh)77.3%Source for GPT-5.3-Codex (xhigh) on Terminal-Bench 2.0 (opens in a new tab)10006
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme