RankingTerminal-Bench 4.0

Terminal-Bench 4.0

Data updated 23 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Terminal-Bench 4.0 reported
Board
https://www.tbench.ai/

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

26–50 of 455 entries

ModelScoreEvidenceSource-recorded date
aya-23-35BCohere
aya-23-8BCohere
aya-expanse-8bCohere
aya-vision-8bCohere
C4Ai Aya Expanse 32BCohere
C4Ai Aya Vision 32BCohere
c4ai-command-r7b-arabic-02-2025Cohere
chatglm-6bZ.ai
chatglm2-6bZ.ai
chatglm2-6b-32kZ.ai
chatglm3-6bZ.ai
chatglm3-6b-128kZ.ai
chatglm3-6b-32kZ.ai
Claude Haiku 4.5Anthropic
Claude Opus 4.5Anthropic
Claude Opus 4.6Anthropic
Claude Opus 4.7Anthropic
Claude Opus 5.5Anthropic
Claude Sonnet 4.5Anthropic
Claude Sonnet 4.6Anthropic
codegeex2-6bZ.ai
codegeex4-all-9bZ.ai
CodeGemma 7B ITGoogle
Codestral 2508Mistral
cogagent-9b-20241220Z.ai