RankingTerminal-Bench 4.0

Terminal-Bench 4.0

Data updated 23 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Terminal-Bench 4.0 reported
Board
https://www.tbench.ai/

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

1–25 of 455 entries

ModelScoreEvidenceSource-recorded date
GPT-6 AstraOpenAI58.18%official board2026-09-03
Claude Fable 5.1Anthropic57.88%official board2026-09-01
Claude Opus 5Anthropic53.94%official board2026-07-24
Claude Fable 5Anthropic44.55%official board2026-06-09
GLM-5.3Z.ai41.82%official board2026-08-14
Grok 4.7xAI37.58%official board2026-09-21
GPT-5.6 SolOpenAI37.27%official board2026-06-26
Claude Opus 4.8Anthropic23.64%official board2026-05-28
GPT-5.6 TerraOpenAI21.52%official board2026-06-26
Grok 4.6xAI20.3%official board2026-08-12
Gemini 3.8 FlashGoogle19.09%official board2026-09-02
GPT-5.6 LunaOpenAI17.27%official board2026-06-26
Claude Sonnet 5Anthropic12.42%official board2026-06-30
Grok-4.5xAI12.42%official board2026-07-16
Gemini 3.7 FlashGoogle11.21%official board2026-08-13
Agnes 2.5 Pro AlphaAgnes AI
Amazon Nova 2 LiteAmazon
Amazon Nova 2 Pro PreviewAmazon
Amazon Nova LiteAmazon
Amazon Nova MicroAmazon
Amazon Nova PremierAmazon
Amazon Nova ProAmazon
AutoGLM-Phone-9BZ.ai
AutoGLM-Phone-9B-MultilingualZ.ai
aya-101Cohere