RankingGemini 3.1 Pro
Gemini 3.1 Pro
86 published benchmark measures · 23 benchmark families contribute across 7 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 6 effort levels across 74 benchmark/harness combinations →
Reported effort · High + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 31 observations
- Not specified: 78 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
23 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Catalog status
- preview
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Gemini 3
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://ai.google.dev/gemini-api/docs/models
- Default Capability family coverage
/badge/gemini-3.1-pro.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agents' Last Exam · not specified | Supporting evidence | 32.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AIME 2026 | Hard reasoning | 98.3%Reported settings & sourceKimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents | Agentic | 32%Reported settings & sourceKimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 33.4%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 77.1% | ARC Prize verifiedcontributes to capability | official board | 2026-02-19 |
| ARC-AGI-3 · 3 | Hard reasoning | 0.42%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.1 · v1.1 | Supporting evidence | 42.7 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index | Supporting evidence | 46 | AA Intelligence Index | official board | 2026-06-15 |
| Artificial Analysis Intelligence Index v4.1 · v4.1 | Supporting evidence | 46.5 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BabyVision | Supporting evidence | 51.6%Reported settings & sourceKimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BabyVision (w/ python) | Supporting evidence | 68.3%Reported settings & sourceKimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BabyVision with tools · source release snapshot; version not specified | Supporting evidence | 51.5%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BankerToolBench · source release snapshot; version not specified | Supporting evidence | 67%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp | Agentic | 85.9% | BrowseComp reported | lab self-report | 2026-02-19 |
| BrowseComp | Agentic | 85.9%Reported settings & sourceKimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BrowseComp · not specified | Agentic | 85.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 85.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp (Agent Swarm) | Agentic | 85.9%Reported settings & sourceKimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) | Multimodal | 80.2%Reported settings & sourceKimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 5 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) (w/ python) | Supporting evidence | 89.9%Reported settings & sourceKimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CharXiv Reasoning with tools · source release snapshot; version not specified | Multimodal | 81.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| CL-bench · source release snapshot; version not specified | Supporting evidence | 21.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Claw Eval (pass@3) | Supporting evidence | 82.9%Reported settings & sourceKimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Claw Eval (pass^3) | Supporting evidence | 57.8%Reported settings & sourceKimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Claw-Eval · source release snapshot; version not specified | Supporting evidence | 57.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA · source release snapshot; version not specified | Agentic | 71.3%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA (accuracy) | Supporting evidence | 60.2%Reported settings & sourceKimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSearchQA (f1-score) | Supporting evidence | 81.9%Reported settings & sourceKimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 11.7% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 11.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 · v1.1 | Coding | 12%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 43%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 1-3 (v2) · v2 | Hard reasoning | 59.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| gdp.pdf · not specified | Supporting evidence | 16.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval rubrics · source release snapshot; version not specified | Agentic | 57.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | 904 | Artificial Analysis GDPval-AAcontributes to capability | official board | 2026-09-12 |
| GDPval-AA v2 · v2 | Agentic | 962Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 Elo · source release snapshot; version not specified | Agentic | 963Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · not specified | Supporting evidence | 3.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 94.141% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 94.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA-Diamond | Supporting evidence | 94.3%Reported settings & sourceKimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional · source release snapshot; version not specified | Supporting evidence | 41.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE with tools · source release snapshot; version not specified | Hard reasoning | 51.4%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE without tools · source release snapshot; version not specified | Hard reasoning | 45.4%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE-Full | Hard reasoning | 44.4%Reported settings & sourceKimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HLE-Full (w/ tools) | Hard reasoning | 51.4%Reported settings & sourceKimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HMMT 2026 (Feb) | Supporting evidence | 94.7%Reported settings & sourceKimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Humanity's Last Exam | Hard reasoning | 46.44% | HLE no toolscontributes to capability | official board | 2026-04-10 |
| IMO 2025 · source release snapshot; version not specified | Supporting evidence | 42.4%Reported settings & sourceMiniMax points out of42; comparator percentages as printed First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| IMO-AnswerBench | Supporting evidence | 91%Reported settings & sourceKimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 15.9%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| KernelBench Hard · source release snapshot; version not specified | Supporting evidence | 18.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LiveCodeBench (v6) | Coding | 91.7%Reported settings & sourceKimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LiveSQLBench · source release snapshot; version not specified | Supporting evidence | 39.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1487 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| Management Consulting Tasks (Internal) · not specified | Supporting evidence | 13.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MathVision | Supporting evidence | 89.8%Reported settings & sourceKimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MathVision (w/ python) | Supporting evidence | 95.7%Reported settings & sourceKimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MCPAtlas · source release snapshot; version not specified | Agentic | 78.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MCPAtlas · source release snapshot; version not specified | Agentic | 69.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MCPMark | Supporting evidence | 55.9%Reported settings & sourceKimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MMLU-Pro | Knowledge | 91.16% | MMLU-Pro reportedcontributes to capability | official board | 2026-09-11 |
| MMMU Pro (no tools) · not specified | Multimodal | 80.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; no tools Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro | Multimodal | 83%Reported settings & sourceKimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 5 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 80.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro (w/ python) | Supporting evidence | 85.3%Reported settings & sourceKimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MRCR v2 1M 8-needle · source release snapshot; version not specified | Long context | 26.3%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 21.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OfficeQA Pro · source release snapshot; version not specified | Agentic | 18.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OJBench (python) | Supporting evidence | 70.7%Reported settings & sourceKimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OmniDocBench · source release snapshot; version not specified | Supporting evidence | 88.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 binary without exec · 2.0 | Agentic | 7.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 partial without exec · 2.0 | Agentic | 30.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| PaperBench · source release snapshot; version not specified | Supporting evidence | 46.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 15.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SciCode | Coding | 58.9%Reported settings & sourceKimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SciCode, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SpreadsheetBench v1 · source release snapshot; version not specified | Supporting evidence | 56.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SVG-Bench · source release snapshot; version not specified | Supporting evidence | 59.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Multilingual | Coding | 76.9%Reported settings & sourceKimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 5 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro | Coding | 46.1% | SWE-bench Pro reportedcontributes to capability | official board | 2026-04-08 |
| SWE-Bench Pro | Coding | 54.2%Reported settings & sourceKimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Bench Pro · not specified | Coding | 54.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 54.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 54.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified | Coding | 80.6% | SWE-bench Verified official | lab self-report | 2026-02-19 |
| SWE-Bench Verified | Coding | 80.6%Reported settings & sourceKimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 80.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-fficiency · source release snapshot; version not specified | Supporting evidence | 19.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWEAtlas-QnA · source release snapshot; version not specified | Supporting evidence | 13.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWEAtlas-TestWriting · source release snapshot; version not specified | Supporting evidence | 29.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.0 (Terminus-2) | Coding | 68.5%Reported settings & sourceKimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 | Coding | 65.62% | Terminal-Bench 2.1 reportedcontributes to capability | official board | 2026-02-19 |
| Terminal-Bench 2.1 · 2.1 | Coding | 70.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 70.3%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 70.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Toolathlon | Agentic | 48.8%Reported settings & sourceKimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · not specified | Agentic | 48.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 61.1%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| USAMO 2026 · source release snapshot; version not specified | Supporting evidence | 74.4%Reported settings & sourceMiniMax points out of42; comparator percentages as printed First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| V* (w/ python) | Supporting evidence | 96.9%Reported settings & sourceKimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| VIBE-V2 · source release snapshot; version not specified | Supporting evidence | 28%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| VideoMME with subtitles · source release snapshot; version not specified | Supporting evidence | 87.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| VideoMMMU · source release snapshot; version not specified | Supporting evidence | 87.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| WebArena Verified · source release snapshot; version not specified | Agentic | 69%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| YC-Bench · source release snapshot; version not specified | Supporting evidence | 1100000 USDReported settings & sourceLaunch figure; agent final assets First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| τ²-bench Telecom | Supporting evidence | 99.3% | τ²-bench Telecom reported | lab self-report | 2026-02-19 |
| LiveCodeBench | Coding | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |