RankingGDPval-AA · 2

GDPval-AA · 2

Data updated 12 Sept 2026

Bucket
Agentic
Unit
elo
Direction
Higher is better
Version
2
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic1853
Reported settings & source

Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Opus 5Anthropic1824
Reported settings & source

Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Fable 5Anthropic1723
Reported settings & source

Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
GPT-5.6 SolOpenAI1711
Reported settings & source

Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Fable 5.1Anthropic1835
Reported settings & source

Artificial Analysis; xhigh effort; same release board snapshot.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.3 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Opus 5Anthropic1824
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Claude Opus 5Anthropic1824
Reported settings & source

Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Claude Sonnet 5Anthropic1584
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.7 FlashGoogle1482
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.8 FlashGoogle1545
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI1710
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI1710
Reported settings & source

Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
GPT-5.6 TerraOpenAI1528
Reported settings & source

Artificial Analysis publicboard snapshot; effort as reported

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Muse Spark 1.2Meta1615
Reported settings & source

Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta1754
Reported settings & source

Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta1709
Reported settings & source

Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06