RankingTerminal-Bench 4.0

Terminal-Bench 4.0

Data updated 23 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Terminal-Bench 4.0 reported
Board
https://www.tbench.ai/

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

126–150 of 455 entries

ModelScoreEvidenceSource-recorded date
Gemini 3.6 FlashGoogle
Gemma 1 2B ITGoogle
Gemma 1 7B ITGoogle
Gemma 1.1 2B ITGoogle
Gemma 1.1 7B ITGoogle
Gemma 2 27B ITGoogle
Gemma 2 2B ITGoogle
Gemma 2 9B ITGoogle
Gemma 3 12B ITGoogle
Gemma 3 1B ITGoogle
Gemma 3 270M ITGoogle
Gemma 3 27B ITGoogle
Gemma 3 4B ITGoogle
Gemma 3n E2B ITGoogle
Gemma 3n E4B ITGoogle
Gemma 4 12B ITGoogle
Gemma 4 26B-A4B ITGoogle
Gemma 4 31B ITGoogle
Gemma 4 E2B ITGoogle
Gemma 4 E4B ITGoogle
GLM-4-32B-0414Z.ai
GLM-4-32B-0414-128KZ.ai
GLM-4-9B-0414Z.ai
glm-4-9b-chatZ.ai
glm-4-9b-chat-1mZ.ai