RankingTerminal-Bench v3.0 · v3.0
Terminal-Bench v3.0 · v3.0
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- v3.0
- Display harness
- Grok 4.6 launch evaluations
- Board
- https://x.ai/news/grok-4-6
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-5.6 SolOpenAI | 34.6%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| Claude Fable 5Anthropic | 34.1%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| Grok 4.6xAI | 26%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| Grok-4.5xAI | 15.7%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |