RankingTerminal-Bench · 4.0

Terminal-Bench · 4.0

Data updated 21 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Grok 4.7 launch comparison
Board
https://x.ai/news/grok-4-7

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–4 of 4 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic57.9%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
Grok 4.7xAI38%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
GPT-5.6 SolOpenAI37.3%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
Grok 4.6xAI20.3%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21