Compare published benchmarks

Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.

These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.

Category weights

Weights are relative. Categories share the total; each benchmark family shares its category equally. Related versions and settings divide their family’s allocation.

GPT-5.6 Sol and Kimi K3

No unambiguous shared results under these settings. Inspect the available results below or choose another source.

Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.

BenchmarkGPT-5.6 SolKimi K3Comparison
ARC-AGI · 1
Hard reasoning · Higher is better

96.5%

Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

No resultMissing result
ARC-AGI · 2
Hard reasoning · Higher is better

92.5%

Reported settings & source

ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

No resultMissing result
BioMysteryBench · Human Solvable
Knowledge · Higher is better

86.1%

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
BioMysteryBench · Human Difficult
Knowledge · Higher is better

28.8%

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
ProteinGym · Hard
Knowledge · Higher is better

35.5 percent rank correlation

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
Organic Chemistry · 2 revised
Knowledge · Higher is better

43.2%

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
Protocols · Troubleshooting
Knowledge · Higher is better

56.4%

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
Protocols · Understanding network-restricted
Knowledge · Higher is better

63.9%

Reported settings & source

Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

No resultMissing result
FrontierSWE · 2
Coding · Higher is better

0.32 fraction

Reported settings & source

Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12

No resultMissing result
Terminal-Bench · 4.0
Agentic · Higher is better

37.3%

Reported settings & source

Claude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board.

Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed.

Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12

No resultMissing result
Terminal-Bench-Science · 0.1
Agentic · Higher is better

22.4%

Reported settings & source

70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12

No resultMissing result
SWE-bench Pro
Coding · Higher is better

64.6%

Reported settings & source

Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

No resultMissing result
CursorBench · 3.2.0
Coding · Higher is better

67.2%

Reported settings & source

Cursor production agent harness; independently measured by Cursor; max effort.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

No resultMissing result
GDPval-AA · 2
Human pref · Higher is better

1711

Reported settings & source

Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

No resultMissing result
AA-Briefcase
Human pref · Higher is better

1502

Reported settings & source

Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12

No resultMissing result
AutomationBench
Agentic · Higher is better

19.6%

Reported settings & source

Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12

No resultMissing result

Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations