RankingARC-AGI · 2

ARC-AGI · 2

Data updated 12 Sept 2026

Bucket
Hard reasoning
Unit
percent
Direction
Higher is better
Version
2
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
GPT-5.6 SolOpenAI92.5%
Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Opus 5Anthropic90.42%
Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Fable 5.1Anthropic90%
Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Fable 5Anthropic89.2%
Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12