RankingMCPAtlas · source release snapshot; version not specified

MCPAtlas · source release snapshot; version not specified

Data updated 12 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
Amazon Nova 2 technical report
Board
https://cdn.amazon.science/c5/3d/84514a224666b5be6de4b43ef4aa/nova-2-0-technical-report2.pdf

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
GPT-5OpenAI44.5%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Sonnet 4.5Anthropic43.8%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova 2 LiteAmazon24.6%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5 MiniOpenAI22.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Gemini 2.5 ProGoogle8.8%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Opus 4.7Anthropic77%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Claude Opus 4.8Anthropic82.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
Claude Sonnet 4.6Anthropic61.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Gemini 3.1 ProGoogle69.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Gemini 3.1 ProGoogle78.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
GLM-5.1Z.ai71.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
GPT-5.5OpenAI75.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
GPT-5.5OpenAI75.3%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06
Kimi-K2.6Moonshot66.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
MiniMax-M2.7MiniMax49.4%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
MiniMax-M3MiniMax74.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

MiniMax M3 model card benchmark figurelab self-report2026-09-06
Muse Spark 1.1Meta88.1%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Muse Spark 1.1 evaluation report Figure44lab self-report2026-09-06