RankingSWE-sweep

SWE-sweep

Data updated 6 Oct 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
2026-09-24
Display harness
SWE-sweep · xhigh
Board
https://swesweep.com/results/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

426–450 of 460 entries

ModelScoreEvidenceSource-recorded date
Seed 2.0 LiteByteDance———
Seed 2.0 MiniByteDance———
Seed 2.0 ProByteDance———
Seed 2.1 ProByteDance———
Seed 2.1 TurboByteDance———
Seed CharacterByteDance———
Seed Coder 8B InstructByteDance———
Seed Coder 8B ReasoningByteDance———
Seed EvolvingByteDance———
Seed OSS 36B InstructByteDance———
SOLAR 10.7B Instruct v1.0Upstage———
Solar Mini (April 2025)Upstage———
Solar Open 100BUpstage———
Solar Open2 250BUpstage———
Solar Pro 2 (December 2025)Upstage———
Solar Pro 3Upstage———
Solar Pro 4Upstage———
solar pro preview instructUpstage———
Step 3StepFun———
Step 3 VL 10BStepFun———
Step 3.5 FlashStepFun———
Step 3.7 FlashStepFun———
SynLogic-32BMiniMax———
SynLogic-7BMiniMax———
SynLogic-Mix-3-32BMiniMax———

Other harnesses

These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.

ModelScoreHarnessEvidenceSource-recorded date
Gemini 3.5 Flash-LiteGoogle0.12%SWE-sweep · unspecifiedofficial board2026-09-24
GPT-5.4 MiniOpenAI0.49%SWE-sweep · highofficial board2026-09-24
GPT-5.4 MiniOpenAI0.2%SWE-sweep · unspecifiedofficial board2026-09-24
GPT-5.6 LunaOpenAI1.4%SWE-sweep · highofficial board2026-09-24
GPT-5.6 LunaOpenAI0.52%SWE-sweep · unspecifiedofficial board2026-09-24
Kimi K3Moonshot0.57%SWE-sweep · unspecifiedofficial board2026-09-24