RankingTerminal-Bench · 4.0
Terminal-Bench · 4.0
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- 4.0
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 55.8%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Opus 5Anthropic | 52.3%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 42%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Opus 5Anthropic | 51.8%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Claude Sonnet 5Anthropic | 12.4%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.7 FlashGoogle | 11.2%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.8 FlashGoogle | 19.1%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 37.3%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 37.3%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 TerraOpenAI | 23.6%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |