RankingGDPval-AA · 2.1
GDPval-AA · 2.1
- Bucket
- Supporting evidence
- Unit
- elo
- Direction
- Higher is better
- Version
- 2.1
- Display harness
- Claude Opus 5.5 System Card
- Board
- https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–2 of 2 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Opus 5.5Anthropic | 1846Reported settings & sourceArtificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; max effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 1846. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.3 · reviewed 2026-09-22 | Claude Opus 5.5 System Card | lab self-report | 2026-09-22 |
| Claude Opus 5.5Anthropic | 1820Reported settings & sourceArtificial Analysis GDPval-AA v2.1; 220 GDPval gold tasks; blind pairwise Elo anchored to DeepSeek V4.1 Flash (max) at 1600; xhigh effort. Run by Artificial Analysis. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.14.3 · reviewed 2026-09-22 | Claude Opus 5.5 System Card | lab self-report | 2026-09-22 |