RankingHealthBench Professional

HealthBench Professional

Data updated 24 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–8 of 8 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic74.2%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic73.4%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic68.9%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic63.3%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5.1Anthropic62.1%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic59.8%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5.5Anthropic77.1%
Reported settings & source

Raw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.15.2 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Opus 5.5Anthropic65.6%
Reported settings & source

Length-adjusted score using the HealthBench Professional paper method; adaptive max; five trials; no tools; Opus 4.8 grader; safety fallback to Opus 5.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's 65.6 is this length-adjusted score, not the raw 77.1%.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.15.2 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22