RankingHealthBench Professional · Length-adjusted
HealthBench Professional · Length-adjusted
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- Length-adjusted
- Display harness
- Claude Haiku 5.5 System Card
- Board
- https://www-cdn.anthropic.com/e1080d6bf5ae2018ea3c2f414064be03232f5be5/Claude%20Haiku%205.5%20System%20Card.pdf
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–5 of 5 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Haiku 5.5Anthropic | 57.9%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Claude Haiku 5.5 System Card | lab self-report | 2026-10-07 |
| Claude Haiku 5.5Anthropic | 59.9%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. medium effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Claude Haiku 5.5 System Card | lab self-report | 2026-10-07 |
| Claude Haiku 5.5Anthropic | 61.3%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. high effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Claude Haiku 5.5 System Card | lab self-report | 2026-10-07 |
| Claude Haiku 5.5Anthropic | 61%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Claude Haiku 5.5 System Card | lab self-report | 2026-10-07 |
| Claude Haiku 5.5Anthropic | 64.8%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Claude Haiku 5.5 System Card | lab self-report | 2026-10-07 |