Skip to main content
Benchmark Registry
ModelsBenchmarksOrganizationsCompare

Terminal-Bench

Benchmark
Terminal-Bench
Evaluated by
Terminal-Bench Team
Release date
November 7, 2025
Version
Terminal-Bench 2.0
Metric
Terminal-Bench Accuracy

Results

6 results

LatestHistory
All providersAnthropicCursorDeepSeek AIGoogle (DeepMind)Meta AI (originally Facebook AI Research)Microsoft AIMiniMaxMistral AIMoonshot AINVIDIAOpenAIZ.ai
Terminal-Bench 2.0 results
ProviderModelScoreSourceRegistry No.
AnthropicClaude Opus 4.7 (no thinking)69.4%Source for Claude Opus 4.7 (no thinking) on Terminal-Bench 2.0 (opens in a new tab)20010
AnthropicClaude Sonnet 4.6 (no thinking, max effort)59.1%Source for Claude Sonnet 4.6 (no thinking, max effort) on Terminal-Bench 2.0 (opens in a new tab)20009
AnthropicClaude Opus 4.6 (adaptive thinking, max)65.4%Source for Claude Opus 4.6 (adaptive thinking, max) on Terminal-Bench 2.0 (opens in a new tab)20008
AnthropicClaude Opus 4.5 (extended thinking (128K))59.8%Source for Claude Opus 4.5 (extended thinking (128K)) on Terminal-Bench 2.0 (opens in a new tab)20007
AnthropicClaude Haiku 4.5 (extended thinking (32K))41.8%Source for Claude Haiku 4.5 (extended thinking (32K)) on Terminal-Bench 2.0 (opens in a new tab)20006
AnthropicClaude Haiku 4.5 (no thinking)40.2%Source for Claude Haiku 4.5 (no thinking) on Terminal-Bench 2.0 (opens in a new tab)20006
Page 1 of 1

© 2026 Densa Labs

Toggle between the data update date and the application build time.
Legal
Color theme