RankingGemini 3.1 Pro

Gemini 3.1 Pro

Data updated 12 Sept 2026

86 published benchmark measures · 23 benchmark families contribute across 7 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Gemini 3.1 Pro capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Gemini 3.1 Pro · Agentic: 24.5 · SupportedGemini 3.1 Pro · Hard reasoning: 35.3 · SupportedGemini 3.1 Pro · Coding: 21.0 · SupportedGemini 3.1 Pro · Human pref: 54.5 · SupportedGemini 3.1 Pro · Multimodal: 21.8 · SupportedGemini 3.1 Pro · Long context: 16.9 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic9 families · 1 with independent evidence · Supported24.5

9 core families; 51 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 22.1–25.4; 0/11 scenarios unsupported. Without one publisher: 20.6–28.7; 0/15 unsupported. Smoothing check: 21.5–30.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 3 with independent evidence · Supported35.3

4 core families; 61 direct opponents across 12 labs.

Disputed order: matched results across at least two families give the opposite order against 2 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1 · Matched comparison 2

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 28.9–42.4; 0/7 scenarios unsupported. Without one publisher: 33.7–40.3; 0/16 unsupported. Smoothing check: 33.2–39.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding5 families · 2 with independent evidence · Supported21.0

5 core families; 33 direct opponents across 10 labs.

Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 19.4–23.3; 0/7 scenarios unsupported. Without one publisher: 19.3–21.9; 0/20 unsupported. Smoothing check: 18.7–27.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported54.5

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 52.1–57.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 95.6 observed win share
    lmarena.ai · Source 1
Knowledge1 families · 1 with independent evidence · PreliminaryUnknown

1 core families; 51 direct opponents across 12 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • MMLU-Pro: 100.0 observed win share
    huggingface.co · Source 1
Multimodal2 families · 0 with independent evidence · Supported21.8

2 core families; 11 direct opponents across 5 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 21.8–21.9; 2/8 scenarios unsupported. Without one publisher: 19.8–21.8; 1/8 unsupported. Smoothing check: 19.0–28.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Long context1 families · 0 with independent evidence · Preliminary16.9

1 core families; 2 direct opponents across 2 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • MRCR: 0.0 observed win share
    Meta · Source 1

Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

Compare 6 effort levels across 74 benchmark/harness combinations →

Reported effort · High + unspecified

Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

Inspect each result and its source ↓ · Download effort evidence

Compare capability profiles →

Score contributions and missing evidence

23 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.

Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

Model information & shareable badge
Lab
Google
Catalog status
preview
Availability
Public provider catalog; account and region restrictions may apply
Family
Gemini 3
Released
Context
License
proprietary
Model card
https://ai.google.dev/gemini-api/docs/models
Default Capability family coverage
Documented-evidence family coverage/badge/gemini-3.1-pro.svg
Benchmark scores & sources

Original results, evaluation harnesses, and evidence behind this model.

BenchmarkBucketScoreHarnessEvidenceSource-recorded date
Agents' Last Exam · not specifiedSupporting evidence32.1%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
AIME 2026Hard reasoning98.3%
Reported settings & source

Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
APEX-AgentsAgentic32%
Reported settings & source

Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
APEX-Agents · source release snapshot; version not specifiedAgentic33.4%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
ARC-AGI-2Hard reasoning77.1%ARC Prize verified

Effort: Not specified

contributes to capability
official board2026-02-19
ARC-AGI-3 · 3Hard reasoning0.42%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence42.7 index score
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Artificial Analysis Intelligence IndexSupporting evidence46 AA Intelligence Index

Effort: Not specified

official board

Supporting evidence outside the reviewed capability core

2026-06-15
Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence46.5 index score
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BabyVisionSupporting evidence51.6%
Reported settings & source

Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
BabyVision (w/ python)Supporting evidence68.3%
Reported settings & source

Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
BabyVision with tools · source release snapshot; version not specifiedSupporting evidence51.5%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BankerToolBench · source release snapshot; version not specifiedSupporting evidence67%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BrowseCompAgentic85.9%BrowseComp reported

Effort: Not specified

lab self-report

Legacy self-report lacks a reviewed comparison configuration

2026-02-19
BrowseCompAgentic85.9%
Reported settings & source

Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
BrowseComp · not specifiedAgentic85.9%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
BrowseComp · source release snapshot; version not specifiedAgentic85.9%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
BrowseComp (Agent Swarm)Agentic85.9%
Reported settings & source

Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
CharXiv (RQ)Multimodal80.2%
Reported settings & source

Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

contributes to capability
lab self-reportReviewed 2026-09-12
CharXiv (RQ) (w/ python)Supporting evidence89.9%
Reported settings & source

Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
CharXiv Reasoning with tools · source release snapshot; version not specifiedMultimodal81.6%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
CL-bench · source release snapshot; version not specifiedSupporting evidence21.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Claw Eval (pass@3)Supporting evidence82.9%
Reported settings & source

Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Claw Eval (pass^3)Supporting evidence57.8%
Reported settings & source

Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Claw-Eval · source release snapshot; version not specifiedSupporting evidence57.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
DeepSearchQA · source release snapshot; version not specifiedAgentic71.3%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
DeepSearchQA (accuracy)Supporting evidence60.2%
Reported settings & source

Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
DeepSearchQA (f1-score)Supporting evidence81.9%
Reported settings & source

Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
DeepSWE v1.1Coding11.7%DeepSWE v1.1 reported

Effort: Not specified

contributes to capability
official board2026-09-03
DeepSWE v1.1 · v1.1Coding11.8%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
DeepSWE v1.1 · v1.1Coding12%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence43%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
FrontierMath Tier 1-3 (v2) · v2Hard reasoning59.6%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
gdp.pdf · not specifiedSupporting evidence16.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
GDPval rubrics · source release snapshot; version not specifiedAgentic57.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GDPval-AAAgentic904Artificial Analysis GDPval-AA

Effort: Not specified

contributes to capability
official board2026-09-12
GDPval-AA v2 · v2Agentic962
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GDPval-AA v2 Elo · source release snapshot; version not specifiedAgentic963
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GeneBench Pro · not specifiedSupporting evidence3.1%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
GPQA DiamondHard reasoning94.141%GPQA Diamond reported

Effort: Not specified

contributes to capability
independent repro2026-09-11
GPQA Diamond · not specifiedHard reasoning94.3%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GPQA-DiamondSupporting evidence94.3%
Reported settings & source

Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
HealthBench Professional · source release snapshot; version not specifiedSupporting evidence41.6%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
HLE with tools · source release snapshot; version not specifiedHard reasoning51.4%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
HLE without tools · source release snapshot; version not specifiedHard reasoning45.4%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
HLE-FullHard reasoning44.4%
Reported settings & source

Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
HLE-Full (w/ tools)Hard reasoning51.4%
Reported settings & source

Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
HMMT 2026 (Feb)Supporting evidence94.7%
Reported settings & source

Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Humanity's Last ExamHard reasoning46.44%HLE no tools

Effort: Not specified

contributes to capability
official board2026-04-10
IMO 2025 · source release snapshot; version not specifiedSupporting evidence42.4%
Reported settings & source

MiniMax points out of42; comparator percentages as printed

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
IMO-AnswerBenchSupporting evidence91%
Reported settings & source

Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
JobBench · source release snapshot; version not specifiedSupporting evidence15.9%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
KernelBench Hard · source release snapshot; version not specifiedSupporting evidence18.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
LiveCodeBench (v6)Coding91.7%
Reported settings & source

Kimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
LiveSQLBench · source release snapshot; version not specifiedSupporting evidence39.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
LMArena Text ArenaHuman pref1487LMArena Text

Effort: Not specified

contributes to capability
official board2026-09-11
Management Consulting Tasks (Internal) · not specifiedSupporting evidence13.2%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
MathVisionSupporting evidence89.8%
Reported settings & source

Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MathVision (w/ python)Supporting evidence95.7%
Reported settings & source

Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MCPAtlas · source release snapshot; version not specifiedAgentic78.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MCPAtlas · source release snapshot; version not specifiedAgentic69.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MCPMarkSupporting evidence55.9%
Reported settings & source

Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MMLU-ProKnowledge91.16%MMLU-Pro reported

Effort: Not specified

contributes to capability
official board2026-09-11
MMMU Pro (no tools) · not specifiedMultimodal80.5%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified; no tools

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MMMU-ProMultimodal83%
Reported settings & source

Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

contributes to capability
lab self-reportReviewed 2026-09-12
MMMU-Pro · source release snapshot; version not specifiedMultimodal80.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MMMU-Pro (w/ python)Supporting evidence85.3%
Reported settings & source

Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MRCR v2 1M 8-needle · source release snapshot; version not specifiedLong context26.3%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
NL2Repo · source release snapshot; version not specifiedSupporting evidence21.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OfficeQA Pro · source release snapshot; version not specifiedAgentic18.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OJBench (python)Supporting evidence70.7%
Reported settings & source

Kimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
OmniDocBench · source release snapshot; version not specifiedSupporting evidence88.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OSWorld 2.0 binary without exec · 2.0Agentic7.8%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OSWorld 2.0 partial without exec · 2.0Agentic30.6%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OSWorld Verified · source release snapshot; version not specifiedAgentic76.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OSWorld Verified · source release snapshot; version not specifiedAgentic76.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
PaperBench · source release snapshot; version not specifiedSupporting evidence46.7%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
PostTrainBench · source release snapshot; version not specifiedSupporting evidence15.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SciCodeCoding58.9%
Reported settings & source

Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SciCode, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
SpreadsheetBench v1 · source release snapshot; version not specifiedSupporting evidence56.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SVG-Bench · source release snapshot; version not specifiedSupporting evidence59.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWE-Bench MultilingualCoding76.9%
Reported settings & source

Kimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

contributes to capability
lab self-reportReviewed 2026-09-12
SWE-bench ProCoding46.1%SWE-bench Pro reported

Effort: Not specified

contributes to capability
official board2026-04-08
SWE-Bench ProCoding54.2%
Reported settings & source

Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
SWE-Bench Pro · not specifiedCoding54.2%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Pro · source release snapshot; version not specifiedCoding54.2%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Pro · source release snapshot; version not specifiedCoding54.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench VerifiedCoding80.6%SWE-bench Verified official

Effort: Not specified

lab self-report

Legacy self-report lacks a reviewed comparison configuration

2026-02-19
SWE-Bench VerifiedCoding80.6%
Reported settings & source

Kimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
SWE-bench Verified · source release snapshot; version not specifiedCoding80.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-fficiency · source release snapshot; version not specifiedSupporting evidence19.7%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWEAtlas-QnA · source release snapshot; version not specifiedSupporting evidence13.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWEAtlas-TestWriting · source release snapshot; version not specifiedSupporting evidence29.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Terminal-Bench 2.0 (Terminus-2)Coding68.5%
Reported settings & source

Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
Terminal-Bench 2.1Coding65.62%Terminal-Bench 2.1 reported

Effort: Not specified

contributes to capability
official board2026-02-19
Terminal-Bench 2.1 · 2.1Coding70.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Terminal-Bench 2.1 · 2.1Coding70.3%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Terminal-Bench 2.1 · 2.1Coding70.3%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

MiniMax cites this comparator from the official Terminal-Bench leaderboard; a common protocol with its internal Terminus 2 run is not established. See MiniMax M3 benchmark figure, Terminal-bench 2.1 methodology.

Reviewed 2026-09-06
ToolathlonAgentic48.8%
Reported settings & source

Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Cited comparator requires original evaluation provenance; not a new matched run.

Reviewed 2026-09-12
Toolathlon · not specifiedAgentic48.8%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / Gemini 3.1 Pro Preview · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Toolathlon Verified · source release snapshot; version not specifiedAgentic61.1%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
USAMO 2026 · source release snapshot; version not specifiedSupporting evidence74.4%
Reported settings & source

MiniMax points out of42; comparator percentages as printed

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
V* (w/ python)Supporting evidence96.9%
Reported settings & source

Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 5 · reviewed 2026-09-12

Published configuration

Effort: High

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
VIBE-V2 · source release snapshot; version not specifiedSupporting evidence28%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
VideoMME with subtitles · source release snapshot; version not specifiedSupporting evidence87.9%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
VideoMMMU · source release snapshot; version not specifiedSupporting evidence87.9%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
WebArena Verified · source release snapshot; version not specifiedAgentic69%
Reported settings & source

Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

First-party reported result; comparator results retain the source evaluation setup.

Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
YC-Bench · source release snapshot; version not specifiedSupporting evidence1100000 USD
Reported settings & source

Launch figure; agent final assets

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
τ²-bench TelecomSupporting evidence99.3%τ²-bench Telecom reported

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

2026-02-19
LiveCodeBenchCoding
OSWorld-VerifiedAgentic