RankingKimi-K2.6

Kimi-K2.6

Data updated 12 Sept 2026

66 published benchmark measures · 19 benchmark families contribute across 7 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Kimi-K2.6 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Kimi-K2.6 · Agentic: 12.8 · SupportedKimi-K2.6 · Hard reasoning: 38.5 · SupportedKimi-K2.6 · Coding: 11.1 · SupportedKimi-K2.6 · Human pref: 33.7 · SupportedKimi-K2.6 · Multimodal: 12.9 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic4 families · 0 with independent evidence · Supported12.8

4 core families; 9 direct opponents across 7 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 9.0–15.6; 0/11 scenarios unsupported. Without one publisher: 11.1–95.8; 0/15 unsupported. Smoothing check: 9.4–20.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 0 with independent evidence · Supported38.5

4 core families; 8 direct opponents across 7 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 33.1–46.6; 0/7 scenarios unsupported. Without one publisher: 16.5–66.7; 1/16 unsupported. Smoothing check: 30.3–45.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding6 families · 0 with independent evidence · Supported11.1

6 core families; 12 direct opponents across 8 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 5.2–38.5; 0/7 scenarios unsupported. Without one publisher: 8.1–34.3; 0/20 unsupported. Smoothing check: 8.5–17.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported33.7

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 25.0–41.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 83.3 observed win share
    lmarena.ai · Source 1
Knowledge1 families · 0 with independent evidence · PreliminaryUnknown

1 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Multimodal2 families · 0 with independent evidence · Preliminary12.9

2 core families; 6 direct opponents across 4 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Long context1 families · 0 with independent evidence · PreliminaryUnknown

1 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Aa Lcr: 100.0 observed win share
    NVIDIA · Source 1

Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

Compare 3 effort levels across 36 benchmark/harness combinations →

Reported effort · Reasoning + unspecified

Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

Inspect each result and its source ↓ · Download effort evidence

Compare capability profiles →

Score contributions and missing evidence

19 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.

Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

Model information & shareable badge
Lab
Moonshot
Catalog status
api-listed
Availability
Documented provider API and open weights for self-hosting
Family
Kimi
Released
Context
256,000 tokens
License
Model card
https://huggingface.co/moonshotai/Kimi-K2.6
Default Capability family coverage
Documented-evidence family coverage/badge/kimi-k2.6.svg
Benchmark scores & sources

Original results, evaluation harnesses, and evidence behind this model.

BenchmarkBucketScoreHarnessEvidenceSource-recorded date
AA-LCR · source release snapshot; version not specifiedLong context70.2%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 70.2

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
AIME 2026Hard reasoning96.4%
Reported settings & source

Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
AIME 2026 · source release snapshot; version not specifiedHard reasoning96.4%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
APEX-AgentsAgentic27.9%
Reported settings & source

Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
Apex-Shortlist (no tools) · source release snapshot; version not specifiedSupporting evidence77.4%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 77.4

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Apex-Shortlist (with tools) · source release snapshot; version not specifiedSupporting evidence73.2%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 73.2

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
BabyVisionSupporting evidence39.8%
Reported settings & source

Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
BabyVision (w/ python)Supporting evidence68.5%
Reported settings & source

Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
BrowseCompAgentic83.2%
Reported settings & source

Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
BrowseComp · source release snapshot; version not specifiedAgentic83.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
BrowseComp · source release snapshot; version not specifiedAgentic61.3%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 61.3

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
BrowseComp (Agent Swarm)Agentic86.3%
Reported settings & source

Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Comparison limit: Agent swarm execution is not an individual model configuration.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Agent swarm execution is not an individual model configuration.

Reviewed 2026-09-12
CharXiv (RQ)Multimodal80.4%
Reported settings & source

Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

contributes to capability
lab self-reportReviewed 2026-09-12
CharXiv (RQ) (w/ python)Supporting evidence86.7%
Reported settings & source

Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Claw Eval (pass@3)Supporting evidence80.9%
Reported settings & source

Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Claw Eval (pass^3)Supporting evidence62.3%
Reported settings & source

Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Claw-Eval · source release snapshot; version not specifiedSupporting evidence61.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
CritPt (no tools) · source release snapshot; version not specifiedSupporting evidence9.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 9.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
DeepSearchQA (accuracy)Supporting evidence83%
Reported settings & source

Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
DeepSearchQA (f1-score)Supporting evidence92.5%
Reported settings & source

Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
GDPVal · source release snapshot; version not specifiedAgentic50.4%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 50.4

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GDPval rubrics · source release snapshot; version not specifiedAgentic65.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GPQA (no tools) · source release snapshot; version not specifiedHard reasoning91%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 91.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GPQA Diamond · source release snapshot; version not specifiedHard reasoning90.5%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
GPQA-DiamondSupporting evidence90.5%
Reported settings & source

Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
HLE (no tools) · source release snapshot; version not specifiedHard reasoning34.8%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 34.8

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
HLE (with tools) · source release snapshot; version not specifiedHard reasoning54%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 54.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
HLE-FullHard reasoning34.7%
Reported settings & source

Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
HLE-Full (w/ tools)Hard reasoning54%
Reported settings & source

Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
HMMT 2026 (Feb)Supporting evidence92.7%
Reported settings & source

Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
HMMT February 2026 · source release snapshot; version not specifiedSupporting evidence92.7%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
IFBench (prompt loose) · source release snapshot; version not specifiedSupporting evidence73.7%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 73.7

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
IMO-AnswerBenchSupporting evidence86%
Reported settings & source

Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
IMOAnswerBench (no tools) · source release snapshot; version not specifiedHard reasoning91.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 91.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
IMOAnswerBench (with tools) · source release snapshot; version not specifiedHard reasoning93.71%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 93.71

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
IOI 2025 · source release snapshot; version not specifiedSupporting evidence585 points
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 585.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
LiveCodeBench (v6)Coding89.6%
Reported settings & source

Kimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
LiveCodeBench (v6) · source release snapshot; version not specifiedCoding90.2%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 90.2

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
LiveCodeBench v6 · source release snapshot; version not specifiedCoding89.6%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
LMArena Text ArenaHuman pref1461LMArena Text

Effort: Not specified

contributes to capability
official board2026-09-11
MathVisionSupporting evidence87.4%
Reported settings & source

Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MathVision (w/ python)Supporting evidence93.2%
Reported settings & source

Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MCPAtlas · source release snapshot; version not specifiedAgentic66.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MCPMarkSupporting evidence55.9%
Reported settings & source

Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
MMLU-Pro · source release snapshot; version not specifiedKnowledge88.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 88.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specifiedKnowledge85%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 85.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MMMU-ProMultimodal79.4%
Reported settings & source

Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

contributes to capability
lab self-reportReviewed 2026-09-12
MMMU-Pro · source release snapshot; version not specifiedMultimodal79.4%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
MMMU-Pro (w/ python)Supporting evidence80.1%
Reported settings & source

Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Multi-Challenge · source release snapshot; version not specifiedSupporting evidence63.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 63.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
NL2Repo · source release snapshot; version not specifiedSupporting evidence42.8%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OJBench (python)Supporting evidence60.6%
Reported settings & source

Kimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
OmniScience Accuracy · source release snapshot; version not specifiedSupporting evidence35.5%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 35.5

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
OSWorld Verified · source release snapshot; version not specifiedAgentic73.1%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
OSWorld-VerifiedAgentic73.1%
Reported settings & source

Kimi K2.6 card, OSWorld-Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, OSWorld-Verified, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
PinchBench · source release snapshot; version not specifiedSupporting evidence90.2%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 90.2

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
ProfBench (Search) · source release snapshot; version not specifiedSupporting evidence56%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 56.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SciCodeCoding52.2%
Reported settings & source

Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SciCode, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
SciCode (subtask) · source release snapshot; version not specifiedCoding52%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 52.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SpreadsheetBench v1 · source release snapshot; version not specifiedSupporting evidence84.5%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SVG-Bench · source release snapshot; version not specifiedSupporting evidence60%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
SWE-Bench MultilingualCoding76.7%
Reported settings & source

Kimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

contributes to capability
lab self-reportReviewed 2026-09-12
SWE-Bench Multilingual · source release snapshot; version not specifiedCoding77.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 77.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-Bench ProCoding58.6%
Reported settings & source

Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
SWE-bench Pro · source release snapshot; version not specifiedCoding58.6%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Pro · source release snapshot; version not specifiedCoding58.6%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-Bench VerifiedCoding80.2%
Reported settings & source

Kimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
SWE-bench Verified · source release snapshot; version not specifiedCoding80.2%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-bench Verified · source release snapshot; version not specifiedCoding80.2%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
SWE-Bench Verified · source release snapshot; version not specifiedCoding75.7%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 75.7

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
TauBench V3 Airline · source release snapshot; version not specifiedSupporting evidence85.8%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 85.8

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
TauBench V3 Average · source release snapshot; version not specifiedSupporting evidence72.4%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 72.4

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
TauBench V3 Banking · source release snapshot; version not specifiedSupporting evidence23.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 23.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
TauBench V3 Retail · source release snapshot; version not specifiedSupporting evidence82.9%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 82.9

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
TauBench V3 Telecom · source release snapshot; version not specifiedSupporting evidence97.8%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 97.8

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Terminal Bench 2.1 · 2.1Coding67.2%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 67.2 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch.

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
Terminal-Bench 2.0 · 2.0Coding66.7%
Reported settings & source

MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

First-party reported result; comparator results retain the source evaluation setup.

MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Microsoft report Table 11 and section 4.1: MAI Terminal-Bench removes timeouts and uses a minimal ReAct harness; comparator values are cited from other model releases. A common evaluation protocol is not established.

Reviewed 2026-09-06
Terminal-Bench 2.0 (Terminus-2)Coding66.7%
Reported settings & source

Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

contributes to capability
lab self-reportReviewed 2026-09-12
Terminal-Bench 2.1 · 2.1Coding53.9%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

contributes to capability
lab self-reportReviewed 2026-09-06
ToolathlonAgentic50%
Reported settings & source

Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

No matched opponent in this evaluation unit

Reviewed 2026-09-12
V* (w/ python)Supporting evidence96.9%
Reported settings & source

Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
Vals.ai Financial Agent 1.1 with web search · 1.1Supporting evidence58.8%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 58.8

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
Vals.ai Financial Agent 1.1 without web search · 1.1Supporting evidence54%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 54.0

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
VIBE-V2 · source release snapshot; version not specifiedSupporting evidence46%
Reported settings & source

MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

First-party reported result; comparator results retain the source evaluation setup.

MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
WideSearch (item-f1)Supporting evidence80.8%
Reported settings & source

Kimi K2.6 card, WideSearch (item-f1). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, WideSearch (item-f1), column 2 · reviewed 2026-09-12

Published configuration

Effort: Reasoning

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-12
WMT24++ (en→xx) · source release snapshot; version not specifiedSupporting evidence84.5 score (source scale)
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 84.5

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06

Published configuration

Effort: Not specified

lab self-report

Supporting evidence outside the reviewed capability core

Reviewed 2026-09-06
ARC-AGI-2Hard reasoning
DeepSWE v1.1Agentic
GDPval-AAAgentic
GPQA DiamondHard reasoning
Humanity's Last ExamHard reasoning
LiveCodeBenchCoding
MMLU-ProKnowledge
OSWorld-VerifiedAgentic
SWE-bench ProAgentic
SWE-bench VerifiedAgentic
Terminal-Bench 2.1Agentic