RankingHealthBench Professional

HealthBench Professional

Data updated 29 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
—
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–8 of 8 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic74.2%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5Anthropic73.4%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5Anthropic68.9%
Reported settings & source

Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5Anthropic63.3%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5.1Anthropic62.1%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5Anthropic59.8%
Reported settings & source

Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5.5Anthropic77.1%
Reported settings & source

Raw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.15.2 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Opus 5.5Anthropic65.6%
Reported settings & source

Length-adjusted score using the HealthBench Professional paper method; adaptive max; five trials; no tools; Opus 4.8 grader; safety fallback to Opus 5.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's 65.6 is this length-adjusted score, not the raw 77.1%.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.15.2 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22