RankingSWE-bench Verified · source release snapshot; version not specified

SWE-bench Verified · source release snapshot; version not specified

Data updated 24 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
MAI-Thinking-1 technical report
Board
https://microsoft.ai/pdf/mai-thinking-1.pdf

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

26–26 of 26 entries

ModelScoreHarnessEvidenceSource-recorded date
Qwen3.5-397B-A17BQwen76.4%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Mistral Medium 3.5 model card performance chartslab self-report2026-09-06