RankingClaude Opus 5.5
Claude Opus 5.5
24 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Auto thinking: 2 observations
- High: 2 observations
- Max: 31 observations
- Medium: 3 observations
- XHigh: 4 observations
- Not specified: 2 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Opus
- Released
- 2026-09-22
- Context
- 1,000,000 tokens
- API list price
- $4 input / $20 output per million tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/models/opus-5-5/overview
- Default Capability family coverage
/badge/claude-opus-5-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Humanity’s Last Exam | Hard reasoning | 64.4%Reported settings & sourceFull 2,500 questions; no tools; section 8.11.1 sets thinking to auto, a 1M total token cap, no compaction, and an Opus 4.6 grader. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's default note says adaptive max effort unless otherwise noted. Section 8.11.1 states thinking was set to auto for these runs. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Humanity’s Last Exam | Hard reasoning | 67.7%Reported settings & sourceFull 2,500 questions; web search, web fetch, programmatic tool calling, and code execution; thinking set to auto; 1M total token cap; no compaction; Opus 4.6 grader; HLE source blocklist and contamination review. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 67.7% with tools. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Legal Agent Benchmark · 120-task held-out subset | Supporting evidence | 8.3%Reported settings & sourceAll-pass rate on Harvey's held-out 120 tasks; max effort. Artificial Analysis harness, as in the previous system card. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Legal Agent Benchmark · 120-task held-out subset | Supporting evidence | 91.2%Reported settings & sourceMean criterion-pass rate on Harvey's held-out 120 tasks; max effort. Artificial Analysis harness, as in the previous system card. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| MILU | Knowledge | 93.1%Reported settings & sourceAverage accuracy across 11 languages; adaptive thinking at max effort; five trials; no tools or custom system prompt. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.16.2; Figure 8.16.2.A · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OfficeQA | Agentic | 78.9%Reported settings & sourceAgentic extracted-text Treasury Bulletin corpus with code execution; max effort; mean of five runs. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OfficeQA Pro | Agentic | 67.7%Reported settings & sourceHarder 133-question subset; agentic extracted-text Treasury Bulletin corpus with code execution; max effort; mean of five runs. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld · 2.0 September 10, 2026 task release | Agentic | 81.8%Reported settings & sourcePartial score, pass@1; 108 tasks; five runs; 1080p; 500 action steps; max reasoning effort; Opus 4.8 grader where required. September 10, 2026 task files and server-side context compaction after 100k tokens. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 81.8% partial. Section 8.13.3 says these settings supersede the Fable 5.1 card's OSWorld 2.0 figures. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.13.3 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld · 2.0 September 10, 2026 task release | Agentic | 48.7%Reported settings & sourceStrict pass rate, pass@1; 108 tasks; five runs; 1080p; 500 action steps; max reasoning effort; Opus 4.8 grader where required. September 10, 2026 task files and server-side context compaction after 100k tokens. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A prints partial/strict as 81.8/48.7. This row is the strict pass rate. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.13.3 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| ProgramBench · 166 golden-task subset | Supporting evidence | 91.2%Reported settings & sourcemini-swe-agent without the six-hour timeout; 34 tasks with a reference binary below 0.9 excluded; scored only on tests the reference binary passes; context up to 1M. Section 8.10.1 does not state reasoning effort. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Hidden-test pass rate, not the fraction of completely solved programs. Effort is not stated in this section. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.10.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| SWE-bench Multilingual | Coding | 93.9%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. 300 problems across nine programming languages. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 93.9%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| SWE-bench Multimodal | Supporting evidence | 61.4%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. Visual context added to issue descriptions. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| SWE-bench Pro | Coding | 89.9%Reported settings & sourceAdaptive thinking at max effort; default sampling; five-trial mean; context at most 1M. SWE-bench Pro problems from actively maintained repositories with large multi-file diffs. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 89.9%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Terminal-Bench · 4.0 | Coding | 66.36%Reported settings & sourceClaude Code --bare; xhigh thinking effort; safeguards enabled with server-side fallback (2.5% of requests, 10% of trials); five trials per task (330 trials) on 66 tasks. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A and the launch grid round this xhigh result to 66.4%. Section 8.5 states 66.36%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Terminal-Bench · 4.0 | Coding | 64.8%Reported settings & sourceClaude Code --bare; max thinking effort; safeguards enabled with the default server-side fallback; five trials per task on 66 tasks. Section 8.5 says this is within noise of xhigh. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.5 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Terminal-Bench-Science · 0.1 | Coding | 58.7%Reported settings & sourceClaude Code --bare; max thinking effort; safeguards enabled with server-side fallback (3.9% of requests, 5% of trials); 10 trials per task (700 trials) on 70 tasks. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. The Anthropic launch grid at https://www.anthropic.com/claude-opus-5-5 shows 58.7%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.6 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Toolathlon · Verified June2026 | Agentic | 77.8%Reported settings & sourcePass@1; 108 tasks; three trials; internal harness mirroring Toolathlon-Verified; adaptive thinking at max effort; safety classifiers on; one safety stop and six sandbox-monitor halts counted as failures. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Toolathlon · Verified June2026 | Agentic | 82.4%Reported settings & sourcePass@3 (at least one of three trials correct); 108 tasks; internal harness; adaptive thinking at max effort; safety classifiers on. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Toolathlon · Verified June2026 | Agentic | 72.2%Reported settings & sourcePass³ (all three trials correct); 108 tasks; internal harness; adaptive thinking at max effort; safety classifiers on. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.14.5.A · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |