RankingSWE-bench Pro
SWE-bench Pro
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- public
- Display harness
- SWE-bench Pro reported
- Board
- https://scale.com/leaderboard/swe_bench_pro_public
Compare published benchmark results with category weights →
Models
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 80% | Anthropic SWE-bench Pro reported system | lab self-report | 2026-09-01 |
| Claude Fable 5.1Anthropic | 81.2% | Anthropic SWE-bench Pro reported system | lab self-report | 2026-09-01 |