RankingMiniMax-M3

MiniMax-M3

Data updated 12 Sept 2026

35 published benchmark measures · 12 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
MiniMax-M3 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100MiniMax-M3 · Agentic: 23.7 · SupportedMiniMax-M3 · Hard reasoning: 55.4 · PreliminaryMiniMax-M3 · Coding: 27.8 · SupportedMiniMax-M3 · Human pref: 21.9 · SupportedMiniMax-M3 · Multimodal: 6.9 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic6 families · 1 with independent evidence · Supported23.7

6 core families; 45 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 17.6–28.5; 0/11 scenarios unsupported. Without one publisher: 20.5–27.1; 1/15 unsupported. Smoothing check: 20.5–30.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 68.6 observed win share
    MiniMax, artificialanalysis.ai · Source 1 Source 2
  • Browsecomp: 71.4 observed win share
    MiniMax · Source 1
  • Apex Agents: 40.0 observed win share
    MiniMax · Source 1
  • OSWorld: 40.0 observed win share
    MiniMax · Source 1
  • OfficeQA: 66.7 observed win share
    MiniMax · Source 1
  • MCP-Atlas: 71.4 observed win share
    MiniMax · Source 1
Hard reasoning1 families · 1 with independent evidence · Preliminary55.4

1 core families; 22 direct opponents across 11 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GPQA: 50.0 observed win share
    artificialanalysis.ai · Source 1
Coding3 families · 0 with independent evidence · Supported27.8

3 core families; 7 direct opponents across 6 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 13.7–45.7; 0/7 scenarios unsupported. Without one publisher: 25.9–43.1; 1/20 unsupported. Smoothing check: 27.8–30.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • SWE-bench Pro: 83.3 observed win share
    MiniMax · Source 1
  • SWE-bench Verified: 50.0 observed win share
    MiniMax · Source 1
  • Terminal-Bench: 100.0 observed win share
    MiniMax · Source 1
Human pref1 families · 1 with independent evidence · Supported21.9

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 11.7–33.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 72.8 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal1 families · 0 with independent evidence · Preliminary6.9

    1 core families; 5 direct opponents across 4 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Mmmu Pro: 40.0 observed win share
      MiniMax · Source 1
    Long contextNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 3 effort levels across 36 benchmark/harness combinations →

      Reported effort · Not specified

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      12 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      MiniMax
      Catalog status
      active
      Availability
      Documented provider API; downloadable official weights
      Family
      MiniMax
      Released
      Context
      1,000,000 tokens
      License
      Model card
      https://platform.minimax.io/docs/api-reference/text-anthropic-api
      Default Capability family coverage
      Documented-evidence family coverage/badge/minimax-m3.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      APEX-Agents · source release snapshot; version not specifiedAgentic27.7%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      BankerToolBench · source release snapshot; version not specifiedSupporting evidence76.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BrowseComp · source release snapshot; version not specifiedAgentic83.5%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      CL-bench · source release snapshot; version not specifiedSupporting evidence20.5%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Claw-Eval · source release snapshot; version not specifiedSupporting evidence74.5%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      DRACO · source release snapshot; version not specifiedSupporting evidence73.2%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      GDPval rubrics · source release snapshot; version not specifiedAgentic74.8%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GDPval-AAAgentic1304Artificial Analysis GDPval-AA

      Effort: Not specified

      contributes to capability
      official board2026-09-12
      GPQA DiamondHard reasoning92.929%GPQA Diamond reported

      Effort: Not specified

      contributes to capability
      independent repro2026-09-11
      IMO 2025 · source release snapshot; version not specifiedSupporting evidence35 points / 42
      Reported settings & source

      MiniMax points out of42; comparator percentages as printed

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      KernelBench Hard · source release snapshot; version not specifiedSupporting evidence28.8%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LiveSQLBench · source release snapshot; version not specifiedSupporting evidence40.2%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LMArena Text ArenaHuman pref1442LMArena Text

      Effort: Not specified

      contributes to capability
      official board2026-09-11
      LOCA-Bench 256k · source release snapshot; version not specifiedSupporting evidence49.3%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      MCPAtlas · source release snapshot; version not specifiedAgentic74.2%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      MMMU-Pro · source release snapshot; version not specifiedMultimodal78.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      NL2Repo · source release snapshot; version not specifiedSupporting evidence42.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      OfficeQA Pro · source release snapshot; version not specifiedAgentic45.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OmniDocBench · source release snapshot; version not specifiedSupporting evidence91.6%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      OSWorld Verified · source release snapshot; version not specifiedAgentic75.2%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      PaperBench · source release snapshot; version not specifiedSupporting evidence52.6%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      PostTrainBench · source release snapshot; version not specifiedSupporting evidence37.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SpreadsheetBench v1 · source release snapshot; version not specifiedSupporting evidence89.4%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SVG-Bench · source release snapshot; version not specifiedSupporting evidence63.7%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SWE-bench Pro · source release snapshot; version not specifiedCoding59%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      SWE-bench Verified · source release snapshot; version not specifiedCoding80.5%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      SWE-fficiency · source release snapshot; version not specifiedSupporting evidence34.8%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SWEAtlas-QnA · source release snapshot; version not specifiedSupporting evidence37.9%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SWEAtlas-TestWriting · source release snapshot; version not specifiedSupporting evidence30.8%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Terminal-Bench 2.1 · 2.1Coding66%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      USAMO 2026 · source release snapshot; version not specifiedSupporting evidence36 points / 42
      Reported settings & source

      MiniMax points out of42; comparator percentages as printed

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      VIBE-V2 · source release snapshot; version not specifiedSupporting evidence50.1%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      VideoMME with subtitles · source release snapshot; version not specifiedSupporting evidence85.4%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      VideoMMMU · source release snapshot; version not specifiedSupporting evidence84.6%
      Reported settings & source

      MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      YC-Bench · source release snapshot; version not specifiedSupporting evidence2100000 USD
      Reported settings & source

      Launch figure; agent final assets

      First-party reported result; comparator results retain the source evaluation setup.

      MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ARC-AGI-2Hard reasoning
      DeepSWE v1.1Agentic
      Humanity's Last ExamHard reasoning
      LiveCodeBenchCoding
      MMLU-ProKnowledge
      OSWorld-VerifiedAgentic
      SWE-bench ProAgentic
      SWE-bench VerifiedAgentic
      Terminal-Bench 2.1Agentic