RankingClaude Opus 5.5

Claude Opus 5.5

Compare models

Data updated 22 Sept 2026

24 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Compare with

Choose a second model and its effort to see both profiles and benchmark differences below.

Compare with

Choose a configuration to compare. The overview below combines evidence across settings and has no model-wide rank.

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Opus 5.5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Mixed settings

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                Anthropic
                Catalog status
                active
                Availability
                Public provider catalog; account and region restrictions may apply
                Family
                Claude Opus
                Released
                2026-09-22
                Context
                1,000,000 tokens
                API list price
                $4 input / $20 output per million tokens
                License
                proprietary
                Model card
                https://platform.claude.com/docs/en/models/opus-5-5/overview
                Default Capability family coverage
                Documented-evidence family coverage/badge/claude-opus-5-5.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                AA-Briefcase · 1.1Supporting evidence1822
                Reported settings & source

                Artificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; max effort. Run by Artificial Analysis.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.4 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AA-Briefcase · 1.1Supporting evidence1780
                Reported settings & source

                Artificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; xhigh effort. Run by Artificial Analysis.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.4 · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AA-Briefcase · 1.1Supporting evidence1705
                Reported settings & source

                Artificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; high effort. Run by Artificial Analysis.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.4 · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBenchAgentic40%
                Reported settings & source

                Zapier private held-out leaderboard; simulated business workflows; every deterministic assertion must pass; max effort.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 40.0%. Launch footnote 2 says Zapier ran these without fallback models and counted safeguard interventions as failures.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.6 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.73 voxel IoU
                Reported settings & source

                Random 1,000 of 17,900 Vision2Code files; five runs; adaptive thinking at max effort; no tools; views rendered at 256x256 px, the resolution Anthropic says matches the reference implementation.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Earlier cards used 128x128 px renders. This section publishes the corrected resolution.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.13.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.962 voxel IoU
                Reported settings & source

                Random 1,000 of 17,900 Vision2Code files; five runs; adaptive thinking at max effort; with tools (container, image files, standard libraries, and an image cropping tool); 256x256 px views.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.13.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                ChartographyMultimodal64.4%
                Reported settings & source

                100 tasks; adaptive thinking at max effort; five runs; no tools; Gemini 3.5 Flash grader; expert acceptable ranges.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.13.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                ChartographyMultimodal89%
                Reported settings & source

                100 tasks; adaptive thinking at max effort; five runs; with tools (container, image file, standard libraries, and an image cropping tool); Gemini 3.5 Flash grader.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 89.0% with tools.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.13.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                CursorBench · 4.0Supporting evidence57.8%
                Reported settings & source

                Cursor production agent harness; max effort. Independently measured by Cursor; Anthropic estimated cost from Cursor token counts.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 57.8%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                CursorBench · 4.0Supporting evidence56%
                Reported settings & source

                Cursor production agent harness; xhigh effort. Independently measured by Cursor.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                CursorBench · 4.0Supporting evidence56%
                Reported settings & source

                Cursor production agent harness; high effort. Independently measured by Cursor.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                CursorBench · 4.0Supporting evidence52.5%
                Reported settings & source

                Cursor production agent harness; medium effort. Independently measured by Cursor.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch-page prose also states 52.5% at default (medium) effort.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE · 1.1Coding74.2%
                Reported settings & source

                113 long-horizon tasks; five-trial mean. Section 8.3 does not state reasoning effort.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Effort is not stated in this section, so it is not copied from the Table 8.1.A max-effort default.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.3 · reviewed 2026-09-22

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode · 1.1 ExtendedSupporting evidence65.3%
                Reported settings & source

                Cognition agentic coding in Claude Code; composite functional and code-quality score; medium effort; mean@5. Highest Extended score. Cognition ran the evaluation.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode · 1.1 ExtendedSupporting evidence63.6%
                Reported settings & source

                Cognition agentic coding in Claude Code; composite functional and code-quality score; max effort; mean@5. Cognition ran the evaluation.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode · 1.1 MainSupporting evidence54.4%
                Reported settings & source

                Cognition agentic coding in Claude Code; composite functional and code-quality score; max effort; mean@5. Cognition ran the evaluation.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 54.4% for FrontierCode v1.1 (Main). Section 8.4 says performance falls above medium effort and this max-effort score is 54.4%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.4 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode · 1.1 MainSupporting evidence54.6%
                Reported settings & source

                Cognition agentic coding in Claude Code; composite functional and code-quality score; medium effort; mean@5. Highest Main score. Cognition ran the evaluation.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch-page prose also states 54.6% at default (medium) effort.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierSWE · 2Supporting evidence62.3%
                Reported settings & source

                Proximal agent harness; max reasoning effort; 34 tasks; five trials per task; mean across trials.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.7 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                GDPval-AA · 2.1Agentic1846
                Reported settings & source

                Artificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; max effort. Run by Artificial Analysis.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 1846.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.3 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                GDPval-AA · 2.1Agentic1820
                Reported settings & source

                Artificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; xhigh effort. Run by Artificial Analysis.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.3 · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                GMMLUKnowledge94.3%
                Reported settings & source

                Average accuracy across 42 languages; adaptive thinking at max effort; single trial; no tools or custom system prompt.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.16.1; Figure 8.16.1.A · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                HealthBenchSupporting evidence68.1%
                Reported settings & source

                Raw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                HealthBenchSupporting evidence60.6%
                Reported settings & source

                Length-adjusted score using the GPT-5.5 system-card method; otherwise the raw HealthBench configuration: adaptive max, five trials, no tools, Opus 4.8 grader, safety fallback to Opus 5.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                HealthBench ProfessionalSupporting evidence77.1%
                Reported settings & source

                Raw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.15.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                HealthBench ProfessionalSupporting evidence65.6%
                Reported settings & source

                Length-adjusted score using the HealthBench Professional paper method; adaptive max; five trials; no tools; Opus 4.8 grader; safety fallback to Opus 5.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's 65.6 is this length-adjusted score, not the raw 77.1%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.15.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22