RankingProgramBench · 166 golden-task subset

ProgramBench · 166 golden-task subset

Data updated 24 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
166 golden-task subset
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–4 of 4 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic87.6%
Reported settings & source

mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

Hidden-test pass rate, not fraction of completely solved programs. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic86.3%
Reported settings & source

mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

Hidden-test pass rate, not fraction of completely solved programs. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic85.4%
Reported settings & source

mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

Hidden-test pass rate, not fraction of completely solved programs.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5.5Anthropic91.2%
Reported settings & source

mini-swe-agent without the six-hour timeout; 34 tasks with a reference binary below 0.9 excluded; scored only on tests the reference binary passes; context up to 1M. Section 8.10.1 does not state reasoning effort.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Hidden-test pass rate, not the fraction of completely solved programs. Effort is not stated in this section.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.10.1 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22