RankingSWE-bench Verified · source release snapshot; version not specified
SWE-bench Verified · source release snapshot; version not specified
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- MAI-Thinking-1 technical report
- Board
- https://microsoft.ai/pdf/mai-thinking-1.pdf
Compare published benchmark results with category weights →
Models
26–26 of 26 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Qwen3.5-397B-A17BQwen | 76.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |