RankingClaude Sonnet 4.6

Claude Sonnet 4.6

Data updated 12 Sept 2026

40 published benchmark measures · 16 benchmark families contribute across 7 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Sonnet 4.6 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Claude Sonnet 4.6 · Agentic: 8.5 · SupportedClaude Sonnet 4.6 · Hard reasoning: 12.6 · SupportedClaude Sonnet 4.6 · Coding: 3.5 · SupportedClaude Sonnet 4.6 · Human pref: 43.0 · SupportedClaude Sonnet 4.6 · Multimodal: 1.2 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic5 families · 1 with independent evidence · Supported8.5

5 core families; 14 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 3.9–12.0; 0/11 scenarios unsupported. Without one publisher: 7.6–38.2; 0/15 unsupported. Smoothing check: 5.5–16.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • OSWorld: 20.0 observed win share
    MiniMax, os-world.github.io · Source 1 Source 2
  • Browsecomp: 25.0 observed win share
    MiniMax, Mistral · Source 1 Source 2
  • Apex Agents: 20.0 observed win share
    MiniMax · Source 1
  • GDPval: 71.4 observed win share
    MiniMax · Source 1
  • MCP-Atlas: 14.3 observed win share
    MiniMax · Source 1
Hard reasoning3 families · 1 with independent evidence · Supported12.6

3 core families; 53 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 9.6–28.1; 0/7 scenarios unsupported. Without one publisher: 8.7–23.6; 0/16 unsupported. Smoothing check: 4.8–25.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 1 with independent evidence · Supported3.5

2 core families; 37 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 1.7–16.8; 2/7 scenarios unsupported. Without one publisher: 2.1–14.8; 1/20 unsupported. Smoothing check: 1.6–9.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported43.0

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 38.5–46.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 89.5 observed win share
    lmarena.ai · Source 1
Knowledge2 families · 0 with independent evidence · PreliminaryUnknown

2 core families; 1 direct opponents across 1 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • MMLU-Pro: 100.0 observed win share
    Microsoft · Source 1
  • Simpleqa Verified: 0.0 observed win share
    Microsoft · Source 1
Multimodal1 families · 0 with independent evidence · Preliminary1.2

1 core families; 5 direct opponents across 5 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Mmmu Pro: 0.0 observed win share
    MiniMax · Source 1
Long context2 families · 0 with independent evidence · PreliminaryUnknown

2 core families; 1 direct opponents across 1 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Graphwalks: 100.0 observed win share
    Microsoft · Source 1
  • Longbench: 100.0 observed win share
    Microsoft · Source 1

Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

Compare 7 effort levels across 68 benchmark/harness combinations →

Reported effort · Maximum reasoning + unspecified

Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

Inspect each result and its source ↓ · Download effort evidence

Compare capability profiles →

Score contributions and missing evidence

16 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.

Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

Model information & shareable badge
Lab
Anthropic
Catalog status
active
Availability
Public provider catalog; account and region restrictions may apply
Family
Claude Sonnet
Released
Context
License
proprietary
Model card
https://platform.claude.com/docs/en/about-claude/model-deprecations
Default Capability family coverage
Documented-evidence family coverage/badge/claude-sonnet-4-6.svg
Benchmark scores & sources

Original results, evaluation harnesses, and evidence behind this model.

BenchmarkBucketScoreHarnessEvidenceSource-recorded date
AdvancedIF rubric-level · source release snapshot; version not specifiedSupporting evidence86%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
AIME 2025 · source release snapshot; version not specifiedHard reasoning95.6%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
AIME 2025 avg@16 · source release snapshot; version not specifiedHard reasoning86.9%
Reported settings & source

Maximum reasoning; Sonnet4.6 external API truncation caveat.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
AllenAI IFBench · source release snapshot; version not specifiedSupporting evidence57.1%
Reported settings & source

Maximum reasoning; Sonnet4.6 external API truncation caveat.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
APEX-Agents · source release snapshot; version not specifiedAgentic26.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
ARC-AGI-2Hard reasoning60.4%ARC Prize verified

Effort: Not specified

contributes to capability
official board2026-02-17
BeyondAIME avg@16 · source release snapshot; version not specifiedSupporting evidence47.3%
Reported settings & source

Maximum reasoning; Sonnet4.6 external API truncation caveat.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BFCL v3 · source release snapshot; version not specifiedSupporting evidence76%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BrowseComp · source release snapshot; version not specifiedAgentic74.7%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
BrowseComp · source release snapshot; version not specifiedAgentic74.7%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Claw-Eval · source release snapshot; version not specifiedSupporting evidence68.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Collie · source release snapshot; version not specifiedSupporting evidence67.7%
Reported settings & source

Maximum reasoning; Sonnet4.6 external API truncation caveat.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
CorpusQA · source release snapshot; version not specifiedSupporting evidence79%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
DeepSWE v1.1Coding29.9%DeepSWE v1.1 reported

Effort: Not specified

contributes to capability
official board2026-09-03
DRACO · source release snapshot; version not specifiedSupporting evidence75.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
GDPval rubrics · source release snapshot; version not specifiedAgentic75.7%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GPQA Diamond · source release snapshot; version not specifiedHard reasoning89.9%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GraphWalks <=128k · source release snapshot; version not specifiedLong context96%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

contributes to capability
lab self-reportReviewed 2026-09-06
HealthBench Professional · source release snapshot; version not specifiedSupporting evidence38%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
IF Bench · source release snapshot; version not specifiedSupporting evidence50%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
LMArena Text ArenaHuman pref1473LMArena Text

Effort: Not specified

contributes to capability
official board2026-09-11
LongBenchV2 · source release snapshot; version not specifiedLong context66%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

contributes to capability
lab self-reportReviewed 2026-09-06
MCPAtlas · source release snapshot; version not specifiedAgentic61.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MedXpertQA · source release snapshot; version not specifiedSupporting evidence49%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
MMLU Pro · source release snapshot; version not specifiedKnowledge87%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

contributes to capability
lab self-reportReviewed 2026-09-06
MMMU-Pro · source release snapshot; version not specifiedMultimodal74.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MultiChallenge · source release snapshot; version not specifiedSupporting evidence57%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OmniDocBench · source release snapshot; version not specifiedSupporting evidence86.9%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OSWorld Verified · source release snapshot; version not specifiedAgentic72.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OSWorld-VerifiedAgentic72.11%OSWorld-Verified reported

Effort: Not specified

contributes to capability
official board2026-03-08
SimpleQA Verified · source release snapshot; version not specifiedKnowledge29%
Reported settings & source

Microsoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06

Published configuration

Effort: Maximum reasoning

contributes to capability
lab self-reportReviewed 2026-09-06
SVG-Bench · source release snapshot; version not specifiedSupporting evidence64.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWE-bench Verified · source release snapshot; version not specifiedCoding79.6%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Verified · source release snapshot; version not specifiedCoding79.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Verified · source release snapshot; version not specifiedCoding79.6%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWEAtlas-QnA · source release snapshot; version not specifiedSupporting evidence31.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWEAtlas-TestWriting · source release snapshot; version not specifiedSupporting evidence31.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
tau3 Airline · source release snapshot; version not specifiedSupporting evidence83%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
tau3 Banking · source release snapshot; version not specifiedSupporting evidence28.4%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
tau3 Retail · source release snapshot; version not specifiedSupporting evidence75.9%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
tau3 Telecom · source release snapshot; version not specifiedSupporting evidence70.4%
Reported settings & source

Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

First-party reported result; comparator results retain the source evaluation setup.

Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Terminal-Bench 2.0 · 2.0Coding59.1%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Microsoft report Table 11 and section 4.1: MAI Terminal-Bench removes timeouts and uses a minimal ReAct harness; comparator values are cited from other model releases. A common evaluation protocol is not established.

Reviewed 2026-09-06
VIBE-V2 · source release snapshot; version not specifiedSupporting evidence42.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
YC-Bench · source release snapshot; version not specifiedSupporting evidence100000 USD
Reported settings & source

Launch figure; agent final assets

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
GDPval-AAAgentic
GPQA DiamondHard reasoning
Humanity's Last ExamHard reasoning
LiveCodeBenchCoding
MMLU-ProKnowledge
SWE-bench ProAgentic
SWE-bench VerifiedAgentic
Terminal-Bench 2.1Agentic