RankingTerminal-Bench · 4.0
Terminal-Bench · 4.0
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- 4.0
- Display harness
- Grok 4.7 launch comparison
- Board
- https://x.ai/news/grok-4-7
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 57.9%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.7xAI | 38%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| GPT-5.6 SolOpenAI | 37.3%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.6xAI | 20.3%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |