RankingTerminal-Bench 4.0

Terminal-Bench 4.0

Data updated 23 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Terminal-Bench 4.0 reported
Board
https://www.tbench.ai/

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

226–250 of 455 entries

ModelScoreEvidenceSource-recorded date
Kimi-Dev-72BMoonshot
Kimi-K2.6Moonshot
Kimi-K2.7-CodeMoonshot
Kimi-Linear-48B-A3B-InstructMoonshot
Kimi-VL-A3B-InstructMoonshot
Kimi-VL-A3B-Thinking-2506Moonshot
Leanstral 1.5Mistral
Llama 4 MaverickMeta
Llama-3.1-405B-InstructMeta
Llama-3.1-70B-InstructMeta
Llama-3.1-8B-InstructMeta
Llama-3.2-11B-Vision-InstructMeta
Llama-3.2-1B-InstructMeta
Llama-3.2-3B-InstructMeta
Llama-3.2-90B-Vision-InstructMeta
Llama-3.3-70B-InstructMeta
Llama-4-Scout-17B-16E-InstructMeta
Magistral Medium 1.2Mistral
Magistral Small 1.2Mistral
MAI-Code-1-FlashMicrosoft
MAI-Code-1.1-FlashMicrosoft
MAI-DS-R1Microsoft
MAI-Thinking-1Microsoft
MiMo-7B-RLXiaomi
MiMo-7B-RL-0530Xiaomi