RankingTerminal-Bench 4.0

Terminal-Bench 4.0

Data updated 23 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Terminal-Bench 4.0 reported
Board
https://www.tbench.ai/

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

301–325 of 455 entries

ModelScoreEvidenceSource-recorded date
NVIDIA-Nemotron-Nano-9B-v2-JapaneseNVIDIA
o1OpenAI
o1-proOpenAI
o3OpenAI
o3-miniOpenAI
o3-proOpenAI
o4-miniOpenAI
Phi-3-medium-128k-instructMicrosoft
Phi-3-medium-4k-instructMicrosoft
Phi-3-mini-128k-instructMicrosoft
Phi-3-mini-4k-instructMicrosoft
Phi-3-small-128k-instructMicrosoft
Phi-3-small-8k-instructMicrosoft
Phi-3-vision-128k-instructMicrosoft
Phi-3.5-mini-instructMicrosoft
Phi-3.5-MoE-instructMicrosoft
Phi-3.5-vision-instructMicrosoft
phi-4Microsoft
Phi-4-mini-flash-reasoningMicrosoft
Phi-4-mini-instructMicrosoft
Phi-4-mini-reasoningMicrosoft
Phi-4-multimodal-instructMicrosoft
Phi-4-reasoningMicrosoft
Phi-4-reasoning-plusMicrosoft
Phi-4-reasoning-vision-15BMicrosoft