Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizationsCompare

Terminal-Bench

Benchmark
Terminal-Bench
Evaluated by
Terminal-Bench Team
Release date
November 7, 2025
Version
Terminal-Bench 2.0
Metric
Terminal-Bench Accuracy

Results

2 results

LatestHistory
All providersAnthropicCursorDeepSeek AIGoogle (DeepMind)Meta AI (originally Facebook AI Research)Microsoft AIMiniMaxMistral AIMoonshot AINVIDIAOpenAIZ.ai
Terminal-Bench 2.0 results
ProviderModelScoreSourceRegistry No.
Google (DeepMind)Gemini 3.1 Pro (High)68.5%Source for Gemini 3.1 Pro (High) on Terminal-Bench 2.0 (opens in a new tab)30005
Google (DeepMind)Gemini 3 Pro (high)54.2%Source for Gemini 3 Pro (high) on Terminal-Bench 2.0 (opens in a new tab)30008
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme