RankingClaude Haiku 5.5
Claude Haiku 5.5
31 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 3 observations
- Low: 4 observations
- Max: 36 observations
- Medium: 5 observations
- Xhigh: 4 observations
- Not specified: 4 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; account and region restrictions may apply.
- Family
- Claude Haiku
- Released
- 2026-10-07
- Context
- 1,000,000 tokens
- API list price
- $0.10 input / $0.50 output per million tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/models/haiku-5-5/overview
- Default Capability family coverage
/badge/claude-haiku-5-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase · 1.1 | Supporting evidence | 1578Reported settings & sourceArtificial Analysis independently ran evaluation; max effort; provider-republished Elo. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary report; no new independent unit. GDPval-AA v2.1 is anchored to DeepSeekV4.1Flash(max)=1600; AA-Briefcase is a different benchmark. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.2–3; p.131 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| AA-Briefcase · 1.1 | Supporting evidence | 1372Reported settings & sourceArtificial Analysis independently ran evaluation; medium effort; provider-republished Elo. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary report; no new independent unit. GDPval-AA v2.1 is anchored to DeepSeekV4.1Flash(max)=1600; AA-Briefcase is a different benchmark. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.2–3; p.131 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| AutomationBench-AA | Supporting evidence | 35.411015132783625% | Artificial Analysis AutomationBench-AA | official board | 2026-10-08 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.67 voxel IoUReported settings & sourceAdaptive thinking at max effort; without tools; mean of five runs. Random1000-file subset of17900 Vision2Code files; execution-verified CadQuery; voxel IoU. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.2; pp.125–127 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.87 voxel IoUReported settings & sourceAdaptive thinking at max effort; with tools; mean of five runs. Random1000-file subset of17900 Vision2Code files; execution-verified CadQuery; voxel IoU. Container with image files, standard libraries and cropping tool. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.2; pp.125–127 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Biomedical image analysis | Supporting evidence | 43.5%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. One attempt/problem; score scaled relative to best public competition results. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.5; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Bug Hunt Bench | Supporting evidence | 20.48% | Bug Hunt Bench · max | official board | 2026-10-07 |
| Bug Hunt Bench | Supporting evidence | 17.43% | Bug Hunt Bench · xhigh | official board | 2026-10-07 |
| Chartography | Multimodal | 46.4%Reported settings & sourceAdaptive thinking at max effort; without tools; mean of five runs. 100 tasks; expert answer ranges; Gemini3.5Flash grader. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.1; pp.123–125 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Chartography | Multimodal | 86.2%Reported settings & sourceAdaptive thinking at max effort; with tools; mean of five runs. 100 tasks; expert answer ranges; Gemini3.5Flash grader. Container with image files, standard libraries and cropping tool. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.9.1; pp.123–125 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| FrontierCode · 1.1 Extended | Supporting evidence | 58.4%Reported settings & sourceCognition evaluation in Claude Code, max effort; composite functional/code-quality score; mean of five runs per task. Main:100 hardest tasks; Extended:150 tasks. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Cognition ran the evaluation; this source republishes the result. Other effort-curve points are not numerically labeled in extracted text and are excluded. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.3; pp.111–114 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| FrontierCode · 1.1 Main | Supporting evidence | 46.4%Reported settings & sourceCognition evaluation in Claude Code, max effort; composite functional/code-quality score; mean of five runs per task. Main:100 hardest tasks; Extended:150 tasks. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Cognition ran the evaluation; this source republishes the result. Other effort-curve points are not numerically labeled in extracted text and are excluded. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.3; pp.111–114 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| FrontierCode · 1.1 Main | Supporting evidence | 45.8%Reported settings & sourceCognition evaluation in Claude Code, xhigh effort; composite functional/code-quality score; mean of five runs per task. Main:100 hardest tasks; Extended:150 tasks. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Cognition ran the evaluation; this source republishes the result. Other effort-curve points are not numerically labeled in extracted text and are excluded. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.3; pp.111–114 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| FrontierSWE · 2 | Supporting evidence | 43.8%Reported settings & sourceProximal agent harness; max reasoning effort; 34 tasks; five trials per task; 20-hour trial budget. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Proximal ran this evaluation; retained as the provider-republished claim. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.6; p.117 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| GDPval-AA · 2.1 | Agentic | 1620Reported settings & sourceArtificial Analysis independently ran evaluation; max effort; provider-republished Elo. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary report; no new independent unit. GDPval-AA v2.1 is anchored to DeepSeekV4.1Flash(max)=1600; AA-Briefcase is a different benchmark. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.2–3; p.131 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| GDPval-AA · 2.1 | Agentic | 1277Reported settings & sourceArtificial Analysis independently ran evaluation; medium effort; provider-republished Elo. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary report; no new independent unit. GDPval-AA v2.1 is anchored to DeepSeekV4.1Flash(max)=1600; AA-Briefcase is a different benchmark. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.10.2–3; p.131 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| GMMLU | Knowledge | 87.8%Reported settings & sourceAdaptive thinking max effort; safety classifiers enabled; single trial; no tools or customized system prompts;42-language mean. Blocked/unanswered0.1% excluded from average. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.12.1; pp.134–135 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench · Length-adjusted | Supporting evidence | 59.9%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort. Length-adjusted score; one run per effort below max onOct5; five-run mean at max onOct2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench · Length-adjusted | Supporting evidence | 60.2%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. medium effort. Length-adjusted score; one run per effort below max onOct5; five-run mean at max onOct2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench · Length-adjusted | Supporting evidence | 60.3%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. high effort. Length-adjusted score; one run per effort below max onOct5; five-run mean at max onOct2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench · Length-adjusted | Supporting evidence | 61.1%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Length-adjusted score; one run per effort below max onOct5; five-run mean at max onOct2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench · Length-adjusted | Supporting evidence | 61.6%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Length-adjusted score; one run per effort below max onOct5; five-run mean at max onOct2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Length-adjusted | Supporting evidence | 57.9%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Length-adjusted | Supporting evidence | 59.9%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. medium effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| HealthBench Professional · Length-adjusted | Supporting evidence | 61.3%Reported settings & sourceSafety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. high effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |