RankingSWE-bench Pro · source release snapshot; version not specified

SWE-bench Pro · source release snapshot; version not specified

Data updated 12 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
Qwen/Qwen3.8-2.4T-A95B
Board
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5Anthropic80%
Reported settings & source

Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

First-party reported result.

Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

Qwen/Qwen3.8-2.4T-A95Blab self-report2026-09-12
Claude Opus 4.8Anthropic69.2%
Reported settings & source

Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

First-party reported result.

Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

Qwen/Qwen3.8-2.4T-A95Blab self-report2026-09-12
Qwen 3.8-MaxQwen67.7%
Reported settings & source

Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

First-party reported result.

Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

Qwen/Qwen3.8-2.4T-A95Blab self-report2026-09-12
GPT-5.6 SolOpenAI64.6%
Reported settings & source

Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

First-party reported result.

Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

Qwen/Qwen3.8-2.4T-A95Blab self-report2026-09-12
Qwen3.7-maxQwen60.6%
Reported settings & source

Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

First-party reported result.

Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

Qwen/Qwen3.8-2.4T-A95Blab self-report2026-09-12
Claude Opus 4.6Anthropic53.4%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

MAI-Thinking-1 technical reportlab self-report2026-09-06
Claude Opus 4.7Anthropic64.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Claude Opus 4.8Anthropic69.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
Gemini 3.1 ProGoogle54.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Gemini 3.1 ProGoogle54.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
GLM-5.1Z.ai58.4%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

MAI-Thinking-1 technical reportlab self-report2026-09-06
GLM-5.1Z.ai58.4%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
GPT-5.4OpenAI57.7%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

MAI-Thinking-1 technical reportlab self-report2026-09-06
GPT-5.5OpenAI58.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
GPT-5.5OpenAI58.6%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
Kimi-K2.6Moonshot58.6%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

MAI-Thinking-1 technical reportlab self-report2026-09-06
Kimi-K2.6Moonshot58.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
MAI-Thinking-1Microsoft52.8%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

MAI-Thinking-1 technical reportlab self-report2026-09-06
MiniMax-M2.7MiniMax56.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
MiniMax-M3MiniMax59%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Muse Spark 1.1Meta61.5%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06