RankingDeepSWE · 1.1

DeepSWE · 1.1

Data updated 12 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
1.1
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic67.4%
Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. 113 tasks; original hidden-test grading.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Sonnet 5Anthropic53.8%
Reported settings & source

Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.7 FlashGoogle65.3%
Reported settings & source

Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.8 FlashGoogle73.7%
Reported settings & source

Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI72.7%
Reported settings & source

Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI73%
Reported settings & source

113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max

Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
GPT-5.6 TerraOpenAI69.6%
Reported settings & source

Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Muse Spark 1.2Meta55%
Reported settings & source

113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning xhigh

Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta75.4%
Reported settings & source

113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max

Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06