RankingClaude Haiku 5.5
Claude Haiku 5.5
31 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 3 observations
- Low: 4 observations
- Max: 36 observations
- Medium: 5 observations
- Xhigh: 4 observations
- Not specified: 4 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; account and region restrictions may apply.
- Family
- Claude Haiku
- Released
- 2026-10-07
- Context
- 1,000,000 tokens
- API list price
- $0.10 input / $0.50 output per million tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/models/haiku-5-5/overview
- Default Capability family coverage
/badge/claude-haiku-5-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| HealthBench Professional · Length-adjusted | Supporting evidence | 61%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Length-adjusted | Supporting evidence | 64.8%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Length-adjusted / Oct4 repeat | Supporting evidence | 64.2%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. Max effort; thirteen repeat runs onOctober4; length-adjusted score, distinct dated rerun. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Raw | Supporting evidence | 61.6%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort; raw score before length adjustment; five-run mean. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Raw | Supporting evidence | 71%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort; raw score before length adjustment; five-run mean. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Humanity’s Last Exam · No tools | Hard reasoning | 45.9%Reported settings & sourceNo tools; max effort from summary table; thinking auto in methodology; 980K task budget, no context compaction; Claude Opus4.6 grader. Reasoning only; no tools. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.8.1; pp.111,118–120 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Humanity’s Last Exam · With tools | Hard reasoning | 57.4%Reported settings & sourceWith tools; max effort from summary table; thinking auto in methodology; 980K task budget, no context compaction; Claude Opus4.6 grader. Web search/fetch, programmatic tool calling, code execution; blocklist and transcript contamination screening. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.8.1; pp.111,118–120 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Medicinal chemistry | Supporting evidence | 58%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. 504 questions about ADME-related measured properties of depicted drug-like compounds. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.3; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| MILU | Knowledge | 87.6%Reported settings & sourceAdaptive thinking max effort; safety classifiers enabled; five-trial mean; no tools or customized system prompts;11-language mean. Blocked/unanswered under0.2% excluded from average. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.12.2; pp.135–136 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Morphology-to-molecule matching | Supporting evidence | 25.4%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Match four human-liver assay image sets to shuffled molecule-dose perturbations. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.2; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| OfficeQA | Agentic | 73.5%Reported settings & sourceAgentic evaluation with relevant Treasury Bulletin documents preselected as extracted text; max effort; five-run mean. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.1; p.130 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| OfficeQA Pro | Agentic | 60.3%Reported settings & sourceAgentic evaluation with relevant Treasury Bulletin documents preselected as extracted text; max effort; five-run mean. Harder133-question subset. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.1; p.130 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| OSWorld · 2.1 offline subset / Partial credit | Agentic | 72.4%Reported settings & sourceOfficial offline82-task subset of108 tasks; max effort; 1080p;500 action steps; five attempts/task; Claude Opus4.8 grader where needed; metric:Partial credit. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.3; pp.127–130 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| OSWorld · 2.1 offline subset / Strict pass rate | Agentic | 37.1%Reported settings & sourceOfficial offline82-task subset of108 tasks; max effort; 1080p;500 action steps; five attempts/task; Claude Opus4.8 grader where needed; metric:Strict pass rate. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.3; pp.127–130 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| PhysicianBench | Supporting evidence | 17.8%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort. Five attempts per task across100 EHR tasks; share of attempts passed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| PhysicianBench | Supporting evidence | 25.2%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. medium effort. Five attempts per task across100 EHR tasks; share of attempts passed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| PhysicianBench | Supporting evidence | 31.6%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. high effort. Five attempts per task across100 EHR tasks; share of attempts passed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| PhysicianBench | Supporting evidence | 35.8%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Five attempts per task across100 EHR tasks; share of attempts passed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| PhysicianBench | Supporting evidence | 43%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Five attempts per task across100 EHR tasks; share of attempts passed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| ProgramBench | Supporting evidence | 82%Reported settings & sourceModified 166-task subset after excluding 34 flaky-reference tasks; hidden-test pass rate; mini-swe-agent harness without upstream six-hour limit; no internet or decompilation. Effort is not explicitly stated in this section. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Not the unmodified 200-task benchmark; scores only tests passed by reference binaries. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.7.1; p.117 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Protein design — Library Ranking | Supporting evidence | 45.3%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Three attempts/problem; prioritize protein designs for experimental testing. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.4; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Protein design — Sequence Generation | Supporting evidence | 33.1%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. One attempt/problem; novel protein sequences conditioned on design constraints. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.4; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Protocol Troubleshooting | Supporting evidence | 58%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Detect and fix molecular-biology protocol issues. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.6; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Protocol Understanding · 2 | Supporting evidence | 64.1%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Benchling real-protocol interpretation/adaptation questions. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.6; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SingleCellBench | Supporting evidence | 56.2%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Bash, file editor and prespecified installed packages. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.1; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |