RankingGDPval-AA · 2
GDPval-AA · 2
- Bucket
- Agentic
- Unit
- elo
- Direction
- Higher is better
- Version
- 2
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
Compare published benchmark results with category weights →
Models
1–20 of 20 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 1853Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Opus 5Anthropic | 1824Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5Anthropic | 1723Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| GPT-5.6 SolOpenAI | 1711Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5.1Anthropic | 1835Reported settings & sourceArtificial Analysis; xhigh effort; same release board snapshot. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.3 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5.1Anthropic | 1735Reported settings & sourceGrok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart. Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence. Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Claude Opus 5Anthropic | 1824Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Claude Opus 5Anthropic | 1824Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| Claude Sonnet 5Anthropic | 1584Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.7 FlashGoogle | 1482Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| Gemini 3.8 FlashGoogle | 1545Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 1710Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 1710Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| GPT-5.6 TerraOpenAI | 1528Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |
| GPT-6 AstraOpenAI | 1542Reported settings & sourceGrok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart. Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence. Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.6xAI | 1605Reported settings & sourceGrok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart. Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence. Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.7xAI | 1695Reported settings & sourceGrok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart. Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence. Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Muse Spark 1.2Meta | 1615Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| Muse Spark 1.3Meta | 1754Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |
| Muse Spark 1.3Meta | 1709Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Muse Spark1.3 evaluation methodology | lab self-report | 2026-09-06 |