RankingTerminal-Bench · 4.0

Terminal-Bench · 4.0

Data updated 24 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
4.0
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–12 of 12 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic55.8%
Reported settings & source

Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic52.3%
Reported settings & source

Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic42%
Reported settings & source

Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic51.8%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Claude Opus 5.5Anthropic66.36%
Reported settings & source

Claude Code --bare; xhigh thinking effort; safeguards enabled with server-side fallback (2.5% of requests, 10% of trials); five trials per task (330 trials) on 66 tasks.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A and the launch grid round this xhigh result to 66.4%. Section 8.5 states 66.36%.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Opus 5.5Anthropic64.8%
Reported settings & source

Claude Code --bare; max thinking effort; safeguards enabled with the default server-side fallback; five trials per task on 66 tasks. Section 8.5 says this is within noise of xhigh.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Sonnet 5Anthropic12.4%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.7 FlashGoogle11.2%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.8 FlashGoogle19.1%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI37.3%
Reported settings & source

Claude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board.

Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed.

Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
GPT-5.6 SolOpenAI37.3%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 TerraOpenAI23.6%
Reported settings & source

Officialpublicboard highest scoring thinking level; nativeagents may differ

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06