RankingSWE-bench Pro · source release snapshot; version not specified
SWE-bench Pro · source release snapshot; version not specified
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- Qwen/Qwen3.8-2.4T-A95B
- Board
- https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 80%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 69.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen 3.8-MaxQwen | 67.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 64.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen3.7-maxQwen | 60.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Opus 4.6Anthropic | 53.4%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | MAI-Thinking-1 technical report | lab self-report | 2026-09-06 |
| Claude Opus 4.7Anthropic | 64.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Claude Opus 4.8Anthropic | 69.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 54.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 54.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 58.4%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | MAI-Thinking-1 technical report | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 58.4%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| GPT-5.4OpenAI | 57.7%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | MAI-Thinking-1 technical report | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 58.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 58.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 58.6%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | MAI-Thinking-1 technical report | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 58.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| MAI-Thinking-1Microsoft | 52.8%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | MAI-Thinking-1 technical report | lab self-report | 2026-09-06 |
| MiniMax-M2.7MiniMax | 56.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| MiniMax-M3MiniMax | 59%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Muse Spark 1.1Meta | 61.5%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |