RankingFrontierSWE · 2

FrontierSWE · 2

Data updated 22 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
2
Display harness
Claude Opus 5.5 System Card
Board
https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–1 of 1 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Opus 5.5Anthropic62.3%
Reported settings & source

Proximal agent harness; max reasoning effort; 34 tasks; five trials per task; mean across trials.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · section 8.7 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22