RankingHealthBench Professional · Raw

HealthBench Professional · Raw

Data updated 8 Oct 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
Raw
Display harness
Claude Haiku 5.5 System Card
Board
https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–2 of 2 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Haiku 5.5Anthropic61.6%
Reported settings & source

Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort; raw score before length adjustment; five-run mean.

Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07

Claude Haiku 5.5 System Cardlab self-report2026-10-07
Claude Haiku 5.5Anthropic71%
Reported settings & source

Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort; raw score before length adjustment; five-run mean.

Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07

Claude Haiku 5.5 System Cardlab self-report2026-10-07