RankingClaude Haiku 5.5

Claude Haiku 5.5

Compare models

Data updated 8 Oct 2026

31 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Haiku 5.5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 13/13 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Mixed settings

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                Anthropic
                Catalog status
                active
                Availability
                Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; account and region restrictions may apply.
                Family
                Claude Haiku
                Released
                2026-10-07
                Context
                1,000,000 tokens
                API list price
                $0.10 input / $0.50 output per million tokens
                Prompts up to 100,000 tokens. Above 100,000 tokens, $0.50 input and $2.50 output per million tokens. Cache reads $0.01/$0.05; 5-minute cache writes $0.125/$0.625 per million tokens. Adaptive thinking defaults to Medium; supported Low, Medium, High, XHigh and Max are not inferred for observations.
                License
                proprietary
                Model card
                https://platform.claude.com/docs/en/models/haiku-5-5/overview
                Default Capability family coverage
                Documented-evidence family coverage/badge/claude-haiku-5-5.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                HealthBench Professional · Length-adjustedSupporting evidence61%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Xhigh

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                HealthBench Professional · Length-adjustedSupporting evidence64.8%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Length-adjusted score; five runs per effort. Low–xhigh onOct4; max onOct1–2.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                HealthBench Professional · Length-adjusted / Oct4 repeatSupporting evidence64.2%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. Max effort; thirteen repeat runs onOctober4; length-adjusted score, distinct dated rerun.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                HealthBench Professional · RawSupporting evidence61.6%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort; raw score before length adjustment; five-run mean.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07

                Published configuration

                Effort: Low

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                HealthBench Professional · RawSupporting evidence71%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort; raw score before length adjustment; five-run mean.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.3; p.133 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Humanity’s Last Exam · No toolsHard reasoning45.9%
                Reported settings & source

                No tools; max effort from summary table; thinking auto in methodology; 980K task budget, no context compaction; Claude Opus4.6 grader. Reasoning only; no tools.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · Table 8.1.A; section 8.8.1; pp.111,118–120 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Humanity’s Last Exam · With toolsHard reasoning57.4%
                Reported settings & source

                With tools; max effort from summary table; thinking auto in methodology; 980K task budget, no context compaction; Claude Opus4.6 grader. Web search/fetch, programmatic tool calling, code execution; blocklist and transcript contamination screening.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · Table 8.1.A; section 8.8.1; pp.111,118–120 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Medicinal chemistrySupporting evidence58%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. 504 questions about ADME-related measured properties of depicted drug-like compounds.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.3; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                MILUKnowledge87.6%
                Reported settings & source

                Adaptive thinking max effort; safety classifiers enabled; five-trial mean; no tools or customized system prompts;11-language mean. Blocked/unanswered under0.2% excluded from average.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.12.2; pp.135–136 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Morphology-to-molecule matchingSupporting evidence25.4%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Match four human-liver assay image sets to shuffled molecule-dose perturbations.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.2; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                OfficeQAAgentic73.5%
                Reported settings & source

                Agentic evaluation with relevant Treasury Bulletin documents preselected as extracted text; max effort; five-run mean.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.10.1; p.130 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                OfficeQA ProAgentic60.3%
                Reported settings & source

                Agentic evaluation with relevant Treasury Bulletin documents preselected as extracted text; max effort; five-run mean. Harder133-question subset.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.10.1; p.130 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                OSWorld · 2.1 offline subset / Partial creditAgentic72.4%
                Reported settings & source

                Official offline82-task subset of108 tasks; max effort; 1080p;500 action steps; five attempts/task; Claude Opus4.8 grader where needed; metric:Partial credit.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.9.3; pp.127–130 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                OSWorld · 2.1 offline subset / Strict pass rateAgentic37.1%
                Reported settings & source

                Official offline82-task subset of108 tasks; max effort; 1080p;500 action steps; five attempts/task; Claude Opus4.8 grader where needed; metric:Strict pass rate.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.9.3; pp.127–130 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                PhysicianBenchSupporting evidence17.8%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. low effort. Five attempts per task across100 EHR tasks; share of attempts passed.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Low

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                PhysicianBenchSupporting evidence25.2%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. medium effort. Five attempts per task across100 EHR tasks; share of attempts passed.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Medium

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                PhysicianBenchSupporting evidence31.6%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. high effort. Five attempts per task across100 EHR tasks; share of attempts passed.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: High

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                PhysicianBenchSupporting evidence35.8%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. xhigh effort. Five attempts per task across100 EHR tasks; share of attempts passed.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Xhigh

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                PhysicianBenchSupporting evidence43%
                Reported settings & source

                Safety classifiers enabled; same grader across models; HealthBench/Professional graded with Claude Opus4.8; PhysicianBench with Claude Opus5. max effort. Five attempts per task across100 EHR tasks; share of attempts passed.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.11.1–3; pp.131–134 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                ProgramBenchSupporting evidence82%
                Reported settings & source

                Modified 166-task subset after excluding 34 flaky-reference tasks; hidden-test pass rate; mini-swe-agent harness without upstream six-hour limit; no internet or decompilation. Effort is not explicitly stated in this section.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Not the unmodified 200-task benchmark; scores only tests passed by reference binaries.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.7.1; p.117 · reviewed 2026-10-07

                Published configuration

                Effort: Not specified

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Protein design — Library RankingSupporting evidence45.3%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Three attempts/problem; prioritize protein designs for experimental testing.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.4; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Protein design — Sequence GenerationSupporting evidence33.1%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. One attempt/problem; novel protein sequences conditioned on design constraints.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.4; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Protocol TroubleshootingSupporting evidence58%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Detect and fix molecular-biology protocol issues.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.6; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                Protocol Understanding · 2Supporting evidence64.1%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Benchling real-protocol interpretation/adaptation questions.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.6; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07
                SingleCellBenchSupporting evidence56.2%
                Reported settings & source

                API, biology safeguards disabled; adaptive thinking max effort; five attempts/problem unless explicitly noted. Scores describe underlying model capability, not deployed safeguards. Bash, file editor and prespecified installed packages.

                Provider-published lab_self_report; retained as launch evidence, not an independent board admission.

                Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Claude Haiku 5.5 System Card · section 8.13.1; pp.136–139 · reviewed 2026-10-07

                Published configuration

                Effort: Max

                lab self-report

                Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

                Reviewed 2026-10-07