RankingClaude Opus 5.5
Claude Opus 5.5
24 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Auto thinking: 2 observations
- High: 2 observations
- Max: 31 observations
- Medium: 3 observations
- XHigh: 4 observations
- Not specified: 2 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Opus
- Released
- 2026-09-22
- Context
- 1,000,000 tokens
- API list price
- $4 input / $20 output per million tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/models/opus-5-5/overview
- Default Capability family coverage
/badge/claude-opus-5-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase · 1.1 | Supporting evidence | 1822Reported settings & sourceArtificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; max effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AA-Briefcase · 1.1 | Supporting evidence | 1780Reported settings & sourceArtificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; xhigh effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AA-Briefcase · 1.1 | Supporting evidence | 1705Reported settings & sourceArtificial Analysis AA-Briefcase v1.1; long-horizon knowledge projects; rubric scoring and pairwise judging; high effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench | Agentic | 40%Reported settings & sourceZapier private held-out leaderboard; simulated business workflows; every deterministic assertion must pass; max effort. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 40.0%. Launch footnote 2 says Zapier ran these without fallback models and counted safeguard interventions as failures. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.6 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.73 voxel IoUReported settings & sourceRandom 1,000 of 17,900 Vision2Code files; five runs; adaptive thinking at max effort; no tools; views rendered at 256x256 px, the resolution Anthropic says matches the reference implementation. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Earlier cards used 128x128 px renders. This section publishes the corrected resolution. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.13.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.962 voxel IoUReported settings & sourceRandom 1,000 of 17,900 Vision2Code files; five runs; adaptive thinking at max effort; with tools (container, image files, standard libraries, and an image cropping tool); 256x256 px views. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.13.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Chartography | Multimodal | 64.4%Reported settings & source100 tasks; adaptive thinking at max effort; five runs; no tools; Gemini 3.5 Flash grader; expert acceptable ranges. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.13.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Chartography | Multimodal | 89%Reported settings & source100 tasks; adaptive thinking at max effort; five runs; with tools (container, image file, standard libraries, and an image cropping tool); Gemini 3.5 Flash grader. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 89.0% with tools. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.13.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| CursorBench · 4.0 | Supporting evidence | 57.8%Reported settings & sourceCursor production agent harness; max effort. Independently measured by Cursor; Anthropic estimated cost from Cursor token counts. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 57.8%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| CursorBench · 4.0 | Supporting evidence | 56%Reported settings & sourceCursor production agent harness; xhigh effort. Independently measured by Cursor. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| CursorBench · 4.0 | Supporting evidence | 56%Reported settings & sourceCursor production agent harness; high effort. Independently measured by Cursor. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| CursorBench · 4.0 | Supporting evidence | 52.5%Reported settings & sourceCursor production agent harness; medium effort. Independently measured by Cursor. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch-page prose also states 52.5% at default (medium) effort. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.8 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE · 1.1 | Coding | 74.2%Reported settings & source113 long-horizon tasks; five-trial mean. Section 8.3 does not state reasoning effort. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Effort is not stated in this section, so it is not copied from the Table 8.1.A max-effort default. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.3 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode · 1.1 Extended | Supporting evidence | 65.3%Reported settings & sourceCognition agentic coding in Claude Code; composite functional and code-quality score; medium effort; mean@5. Highest Extended score. Cognition ran the evaluation. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode · 1.1 Extended | Supporting evidence | 63.6%Reported settings & sourceCognition agentic coding in Claude Code; composite functional and code-quality score; max effort; mean@5. Cognition ran the evaluation. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode · 1.1 Main | Supporting evidence | 54.4%Reported settings & sourceCognition agentic coding in Claude Code; composite functional and code-quality score; max effort; mean@5. Cognition ran the evaluation. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 54.4% for FrontierCode v1.1 (Main). Section 8.4 says performance falls above medium effort and this max-effort score is 54.4%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode · 1.1 Main | Supporting evidence | 54.6%Reported settings & sourceCognition agentic coding in Claude Code; composite functional and code-quality score; medium effort; mean@5. Highest Main score. Cognition ran the evaluation. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch-page prose also states 54.6% at default (medium) effort. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.4 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierSWE · 2 | Supporting evidence | 62.3%Reported settings & sourceProximal agent harness; max reasoning effort; 34 tasks; five trials per task; mean across trials. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.7 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| GDPval-AA · 2.1 | Agentic | 1846Reported settings & sourceArtificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; max effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 1846. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.3 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| GDPval-AA · 2.1 | Agentic | 1820Reported settings & sourceArtificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; xhigh effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.3 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| GMMLU | Knowledge | 94.3%Reported settings & sourceAverage accuracy across 42 languages; adaptive thinking at max effort; single trial; no tools or custom system prompt. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.16.1; Figure 8.16.1.A · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| HealthBench | Supporting evidence | 68.1%Reported settings & sourceRaw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| HealthBench | Supporting evidence | 60.6%Reported settings & sourceLength-adjusted score using the GPT-5.5 system-card method; otherwise the raw HealthBench configuration: adaptive max, five trials, no tools, Opus 4.8 grader, safety fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| HealthBench Professional | Supporting evidence | 77.1%Reported settings & sourceRaw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.15.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| HealthBench Professional | Supporting evidence | 65.6%Reported settings & sourceLength-adjusted score using the HealthBench Professional paper method; adaptive max; five trials; no tools; Opus 4.8 grader; safety fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's 65.6 is this length-adjusted score, not the raw 77.1%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.15.2 · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |