RankingSWE-bench Pro
SWE-bench Pro
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- —
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 81.2%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 80%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Opus 5Anthropic | 79.2%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 64.6%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |