RankingDeepSWE v1.1
DeepSWE v1.1
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- 1.1
- Display harness
- DeepSWE v1.1 reported
- Board
- https://deepswe.datacurve.ai/
Compare published benchmark results with category weights →
Models
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 67.4% | Anthropic DeepSWE1.1 reported system | lab self-report | 2026-09-01 |