RankingOSWorld Verified · source release snapshot; version not specified
OSWorld Verified · source release snapshot; version not specified
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- MiniMax M3 model card benchmark figure
- Board
- https://huggingface.co/MiniMaxAI/MiniMax-M3
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Opus 4.7Anthropic | 82.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 78.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 76.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| MiniMax-M3MiniMax | 75.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 73.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Claude Sonnet 4.6Anthropic | 72.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Claude Opus 4.8Anthropic | 83.4%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 76.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 78.7%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Muse Spark 1.1Meta | 80.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |