RankingDeepSWE · 1.1
DeepSWE · 1.1
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- 1.1
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 67.4%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. 113 tasks; original hidden-test grading. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.3 · reviewed 2026-09-12 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-12 |
| Claude Sonnet 5Anthropic | 53.8%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.7 FlashGoogle | 65.3%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.8 FlashGoogle | 73.7%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 72.7%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 73%Reported settings & source113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| GPT-5.6 TerraOpenAI | 69.6%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Muse Spark 1.2Meta | 55%Reported settings & source113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning xhigh Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| Muse Spark 1.3Meta | 75.4%Reported settings & source113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |