RankingClaude Haiku 5.5
Claude Haiku 5.5
31 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 3 observations
- Low: 4 observations
- Max: 36 observations
- Medium: 5 observations
- Xhigh: 4 observations
- Not specified: 4 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; account and region restrictions may apply.
- Family
- Claude Haiku
- Released
- 2026-10-07
- Context
- 1,000,000 tokens
- API list price
- $0.10 input / $0.50 output per million tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/models/haiku-5-5/overview
- Default Capability family coverage
/badge/claude-haiku-5-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| SpatialBench · Verified | Supporting evidence | 67.7%Reported settings & sourceAPI, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Bash, file editor and prespecified installed packages. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.13.1; pp.136–139 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SWE-bench Multilingual | Coding | 83.7%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; evaluation-dependent context at most 1M (Table 8.1.A). 300 problems across nine programming languages. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.2; pp.111–112 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SWE-bench Multimodal | Supporting evidence | 30.7%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; evaluation-dependent context at most 1M (Table 8.1.A). Visual context added to issue descriptions. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.2; pp.111–112 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SWE-bench Pro | Coding | 64.8%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; evaluation-dependent context at most 1M (Table 8.1.A). Actively maintained repositories; large multi-file diffs. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · Table 8.1.A; section 8.2; pp.111–112 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Terminal-Bench · 4.0 | Coding | 39.2%Reported settings & sourceClaude Code --bare; max thinking effort; 66 tasks, ten trials per task (660 trials); no internet egress; safeguards enabled, no fallback. Blocked requests ended the trial (1.8% of trials, all failed). Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.4; pp.115–116 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Terminal-Bench-Science · 0.1 | Coding | 20.6%Reported settings & sourceClaude Code --bare; max thinking effort; 70 tasks, 700 trials; no internet egress; safeguards enabled, no fallback; 2.3% of trials blocked and failed. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Claude Haiku 5.5 System Card · section 8.5; p.116 · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |