RankingClaude Opus 5.5

Claude Opus 5.5

Compare models

Data updated 22 Sept 2026

24 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Opus 5.5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Mixed settings

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                Anthropic
                Catalog status
                active
                Availability
                Public provider catalog; account and region restrictions may apply
                Family
                Claude Opus
                Released
                2026-09-22
                Context
                1,000,000 tokens
                API list price
                $4 input / $20 output per million tokens
                License
                proprietary
                Model card
                https://platform.claude.com/docs/en/models/opus-5-5/overview
                Default Capability family coverage
                Documented-evidence family coverage/badge/claude-opus-5-5.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                Humanity’s Last ExamHard reasoning64.4%
                Reported settings & source

                Full 2,500 questions; no tools; section 8.11.1 sets thinking to auto, a 1M total token cap, no compaction, and an Opus 4.6 grader.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's default note says adaptive max effort unless otherwise noted. Section 8.11.1 states thinking was set to auto for these runs.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

                Published configuration

                Effort: Auto thinking

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Humanity’s Last ExamHard reasoning67.7%
                Reported settings & source

                Full 2,500 questions; web search, web fetch, programmatic tool calling, and code execution; thinking set to auto; 1M total token cap; no compaction; Opus 4.6 grader; HLE source blocklist and contamination review.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 67.7% with tools.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

                Published configuration

                Effort: Auto thinking

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Legal Agent Benchmark · 120-task held-out subsetSupporting evidence8.3%
                Reported settings & source

                All-pass rate on Harvey's held-out 120 tasks; max effort. Artificial Analysis harness, as in the previous system card.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Legal Agent Benchmark · 120-task held-out subsetSupporting evidence91.2%
                Reported settings & source

                Mean criterion-pass rate on Harvey's held-out 120 tasks; max effort. Artificial Analysis harness, as in the previous system card.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                MILUKnowledge93.1%
                Reported settings & source

                Average accuracy across 11 languages; adaptive thinking at max effort; five trials; no tools or custom system prompt.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.16.2; Figure 8.16.2.A · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OfficeQAAgentic78.9%
                Reported settings & source

                Agentic extracted-text Treasury Bulletin corpus with code execution; max effort; mean of five runs.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OfficeQA ProAgentic67.7%
                Reported settings & source

                Harder 133-question subset; agentic extracted-text Treasury Bulletin corpus with code execution; max effort; mean of five runs.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.14.1 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld · 2.0 September 10, 2026 task releaseAgentic81.8%
                Reported settings & source

                Partial score, pass@1; 108 tasks; five runs; 1080p; 500 action steps; max reasoning effort; Opus 4.8 grader where required. September 10, 2026 task files and server-side context compaction after 100k tokens.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 81.8% partial. Section 8.13.3 says these settings supersede the Fable 5.1 card's OSWorld 2.0 figures.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.13.3 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld · 2.0 September 10, 2026 task releaseAgentic48.7%
                Reported settings & source

                Strict pass rate, pass@1; 108 tasks; five runs; 1080p; 500 action steps; max reasoning effort; Opus 4.8 grader where required. September 10, 2026 task files and server-side context compaction after 100k tokens.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A prints partial/strict as 81.8/48.7. This row is the strict pass rate.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.13.3 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                ProgramBench · 166 golden-task subsetSupporting evidence91.2%
                Reported settings & source

                mini-swe-agent without the six-hour timeout; 34 tasks with a reference binary below 0.9 excluded; scored only on tests the reference binary passes; context up to 1M. Section 8.10.1 does not state reasoning effort.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Hidden-test pass rate, not the fraction of completely solved programs. Effort is not stated in this section.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.10.1 · reviewed 2026-09-22

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                SWE-bench MultilingualCoding93.9%
                Reported settings & source

                Adaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. 300 problems across nine programming languages.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 93.9%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                SWE-bench MultimodalSupporting evidence61.4%
                Reported settings & source

                Adaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. Visual context added to issue descriptions.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                SWE-bench ProCoding89.9%
                Reported settings & source

                Adaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. SWE-bench Pro problems from actively maintained repositories with large multi-file diffs.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 89.9%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Terminal-Bench · 4.0Coding66.36%
                Reported settings & source

                Claude Code --bare; xhigh thinking effort; safeguards enabled with server-side fallback (2.5% of requests, 10% of trials); five trials per task (330 trials) on 66 tasks.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A and the launch grid round this xhigh result to 66.4%. Section 8.5 states 66.36%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Terminal-Bench · 4.0Coding64.8%
                Reported settings & source

                Claude Code --bare; max thinking effort; safeguards enabled with the default server-side fallback; five trials per task on 66 tasks. Section 8.5 says this is within noise of xhigh.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Terminal-Bench-Science · 0.1Coding58.7%
                Reported settings & source

                Claude Code --bare; max thinking effort; safeguards enabled with server-side fallback (3.9% of requests, 5% of trials); 10 trials per task (700 trials) on 70 tasks.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. The Anthropic launch grid at https://www.anthropic.com/claude-opus-5-5 shows 58.7%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.1.A; section 8.6 · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Toolathlon · Verified June2026Agentic77.8%
                Reported settings & source

                Pass@1; 108 tasks; three trials; internal harness mirroring Toolathlon-Verified; adaptive thinking at max effort; safety classifiers on; one safety stop and six sandbox-monitor halts counted as failures.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Toolathlon · Verified June2026Agentic82.4%
                Reported settings & source

                Pass@3 (at least one of three trials correct); 108 tasks; internal harness; adaptive thinking at max effort; safety classifiers on.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Toolathlon · Verified June2026Agentic72.2%
                Reported settings & source

                Pass³ (all three trials correct); 108 tasks; internal harness; adaptive thinking at max effort; safety classifiers on.

                Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                ARC-AGI-2Hard reasoning
                DeepSWE v1.1Agentic
                GDPval-AAAgentic
                GPQA DiamondHard reasoning
                Humanity's Last ExamHard reasoning
                LiveCodeBenchCoding