RankingSWE-bench Pro

SWE-bench Pro

Data updated 24 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–5 of 5 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic81.2%
Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic80%
Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic79.2%
Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
GPT-5.6 SolOpenAI64.6%
Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5.5Anthropic89.9%
Reported settings & source

Adaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. SWE-bench Pro problems from actively maintained repositories with large multi-file diffs.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 89.9%.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22