Compare published benchmarks
Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.
These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.
GPT-5.6 Sol and Kimi K3
No unambiguous shared results under these settings. Inspect the available results below or choose another source.
Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.
| Benchmark | GPT-5.6 Sol | Kimi K3 | Comparison |
|---|---|---|---|
| ARC-AGI · 1 | 96.5% ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | No result | Missing result |
| ARC-AGI · 2 | 92.5% ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | No result | Missing result |
| BioMysteryBench · Human Solvable | 86.1% Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| BioMysteryBench · Human Difficult | 28.8% Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| ProteinGym · Hard | 35.5 percent rank correlation Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| Organic Chemistry · 2 revised | 43.2% Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| Protocols · Troubleshooting | 56.4% Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| Protocols · Understanding network-restricted | 63.9% Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | No result | Missing result |
| FrontierSWE · 2 | 0.32 fraction Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12 | No result | Missing result |
| Terminal-Bench · 4.0 | 37.3% Claude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | No result | Missing result |
| Terminal-Bench-Science · 0.1 | 22.4% 70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12 | No result | Missing result |
| SWE-bench Pro | 64.6% Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | No result | Missing result |
| CursorBench · 3.2.0 | 67.2% Cursor production agent harness; independently measured by Cursor; max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12 | No result | Missing result |
| GDPval-AA · 2 | 1711 Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12 | No result | Missing result |
| AA-Briefcase | 1502 Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12 | No result | Missing result |
| AutomationBench | 19.6% Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12 | No result | Missing result |
Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations