RankingReal-SWE
Real-SWE
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- 2026-09
- Display harness
- Real-SWE · Fable 5.1 · Claude Code
- Board
- https://realswe.withspecific.com/
Real-SWE ranks each model with its native harness. These rows are display only. They do not enter the weighted Overall or capability scores.
Compare published benchmark results with category weights →
Models
76–100 of 455 entries
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Gemini 3.8 FlashGoogle | 31.2% | Real-SWE · Gemini 3.8 Flash · Gemini CLI | official board | 2026-09-23 |
| GLM-5.3Z.ai | 28.8% | Real-SWE · GLM 5.3 · Claude Code | official board | 2026-09-23 |
| GPT-5.6 SolOpenAI | 16.2% | Real-SWE · GPT-5.6 Sol · Codex CLI | official board | 2026-09-23 |
| GPT-6 AstraOpenAI | 33.8% | Real-SWE · GPT-6 Astra · Codex CLI | official board | 2026-09-23 |
| Grok 4.6xAI | 32.5% | Real-SWE · Grok 4.6 · Grok Build | official board | 2026-09-23 |
| Kimi K3Moonshot | 18.8% | Real-SWE · Kimi K3 · Kimi Code | official board | 2026-09-23 |
| Muse Spark 1.3Meta | 23.8% | Real-SWE · Muse Spark 1.3 · Muse Code | official board | 2026-09-23 |