RankingSWE-sweep
SWE-sweep
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- 2026-09-24
- Display harness
- SWE-sweep · xhigh
- Board
- https://swesweep.com/results/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
76–100 of 460 entries
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Gemini 3.5 Flash-LiteGoogle | 0.12% | SWE-sweep · unspecified | official board | 2026-09-24 |
| GPT-5.4 MiniOpenAI | 0.49% | SWE-sweep · high | official board | 2026-09-24 |
| GPT-5.4 MiniOpenAI | 0.2% | SWE-sweep · unspecified | official board | 2026-09-24 |
| GPT-5.6 LunaOpenAI | 1.4% | SWE-sweep · high | official board | 2026-09-24 |
| GPT-5.6 LunaOpenAI | 0.52% | SWE-sweep · unspecified | official board | 2026-09-24 |
| Kimi K3Moonshot | 0.57% | SWE-sweep · unspecified | official board | 2026-09-24 |