RankingKimi-K2.6
Kimi-K2.6
66 published benchmark measures · 19 benchmark families contribute across 7 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 3 effort levels across 36 benchmark/harness combinations →
Reported effort · Reasoning + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Reasoning: 33 observations
- Not specified: 53 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
19 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Moonshot
- Catalog status
- api-listed
- Availability
- Documented provider API and open weights for self-hosting
- Family
- Kimi
- Released
- —
- Context
- 256,000 tokens
- License
- —
- Model card
- https://huggingface.co/moonshotai/Kimi-K2.6
- Default Capability family coverage
/badge/kimi-k2.6.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-LCR · source release snapshot; version not specified | Long context | 70.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 70.2 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AIME 2026 | Hard reasoning | 96.4%Reported settings & sourceKimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AIME 2026 · source release snapshot; version not specified | Hard reasoning | 96.4%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| APEX-Agents | Agentic | 27.9%Reported settings & sourceKimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Apex-Shortlist (no tools) · source release snapshot; version not specified | Supporting evidence | 77.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 77.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Apex-Shortlist (with tools) · source release snapshot; version not specified | Supporting evidence | 73.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 73.2 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BabyVision | Supporting evidence | 39.8%Reported settings & sourceKimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BabyVision (w/ python) | Supporting evidence | 68.5%Reported settings & sourceKimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BrowseComp | Agentic | 83.2%Reported settings & sourceKimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BrowseComp · source release snapshot; version not specified | Agentic | 83.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 61.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 61.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp (Agent Swarm) | Agentic | 86.3%Reported settings & sourceKimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Comparison limit: Agent swarm execution is not an individual model configuration. Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) | Multimodal | 80.4%Reported settings & sourceKimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) (w/ python) | Supporting evidence | 86.7%Reported settings & sourceKimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Claw Eval (pass@3) | Supporting evidence | 80.9%Reported settings & sourceKimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Claw Eval (pass^3) | Supporting evidence | 62.3%Reported settings & sourceKimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Claw-Eval · source release snapshot; version not specified | Supporting evidence | 61.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CritPt (no tools) · source release snapshot; version not specified | Supporting evidence | 9.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 9.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA (accuracy) | Supporting evidence | 83%Reported settings & sourceKimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSearchQA (f1-score) | Supporting evidence | 92.5%Reported settings & sourceKimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPVal · source release snapshot; version not specified | Agentic | 50.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 50.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval rubrics · source release snapshot; version not specified | Agentic | 65.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA (no tools) · source release snapshot; version not specified | Hard reasoning | 91%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 91.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 90.5%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA-Diamond | Supporting evidence | 90.5%Reported settings & sourceKimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HLE (no tools) · source release snapshot; version not specified | Hard reasoning | 34.8%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 34.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE (with tools) · source release snapshot; version not specified | Hard reasoning | 54%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 54.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE-Full | Hard reasoning | 34.7%Reported settings & sourceKimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HLE-Full (w/ tools) | Hard reasoning | 54%Reported settings & sourceKimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HMMT 2026 (Feb) | Supporting evidence | 92.7%Reported settings & sourceKimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HMMT February 2026 · source release snapshot; version not specified | Supporting evidence | 92.7%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| IFBench (prompt loose) · source release snapshot; version not specified | Supporting evidence | 73.7%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 73.7 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| IMO-AnswerBench | Supporting evidence | 86%Reported settings & sourceKimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| IMOAnswerBench (no tools) · source release snapshot; version not specified | Hard reasoning | 91.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 91.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IMOAnswerBench (with tools) · source release snapshot; version not specified | Hard reasoning | 93.71%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 93.71 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IOI 2025 · source release snapshot; version not specified | Supporting evidence | 585 pointsReported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 585.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LiveCodeBench (v6) | Coding | 89.6%Reported settings & sourceKimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LiveCodeBench (v6) · source release snapshot; version not specified | Coding | 90.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 90.2 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| LiveCodeBench v6 · source release snapshot; version not specified | Coding | 89.6%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1461 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| MathVision | Supporting evidence | 87.4%Reported settings & sourceKimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MathVision (w/ python) | Supporting evidence | 93.2%Reported settings & sourceKimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MCPAtlas · source release snapshot; version not specified | Agentic | 66.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MCPMark | Supporting evidence | 55.9%Reported settings & sourceKimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MMLU-Pro · source release snapshot; version not specified | Knowledge | 88.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 88.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specified | Knowledge | 85%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 85.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro | Multimodal | 79.4%Reported settings & sourceKimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 79.4%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro (w/ python) | Supporting evidence | 80.1%Reported settings & sourceKimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Multi-Challenge · source release snapshot; version not specified | Supporting evidence | 63.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 63.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 42.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OJBench (python) | Supporting evidence | 60.6%Reported settings & sourceKimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OmniScience Accuracy · source release snapshot; version not specified | Supporting evidence | 35.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 35.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld Verified · source release snapshot; version not specified | Agentic | 73.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld-Verified | Agentic | 73.1%Reported settings & sourceKimi K2.6 card, OSWorld-Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, OSWorld-Verified, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PinchBench · source release snapshot; version not specified | Supporting evidence | 90.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 90.2 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ProfBench (Search) · source release snapshot; version not specified | Supporting evidence | 56%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 56.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SciCode | Coding | 52.2%Reported settings & sourceKimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SciCode, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SciCode (subtask) · source release snapshot; version not specified | Coding | 52%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 52.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SpreadsheetBench v1 · source release snapshot; version not specified | Supporting evidence | 84.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SVG-Bench · source release snapshot; version not specified | Supporting evidence | 60%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Multilingual | Coding | 76.7%Reported settings & sourceKimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-Bench Multilingual · source release snapshot; version not specified | Coding | 77.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 77.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Pro | Coding | 58.6%Reported settings & sourceKimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 58.6%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 58.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Verified | Coding | 80.2%Reported settings & sourceKimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 80.2%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 80.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Verified · source release snapshot; version not specified | Coding | 75.7%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 75.7 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Airline · source release snapshot; version not specified | Supporting evidence | 85.8%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 85.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Average · source release snapshot; version not specified | Supporting evidence | 72.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 72.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Banking · source release snapshot; version not specified | Supporting evidence | 23.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 23.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Retail · source release snapshot; version not specified | Supporting evidence | 82.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 82.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Telecom · source release snapshot; version not specified | Supporting evidence | 97.8%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 97.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal Bench 2.1 · 2.1 | Coding | 67.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 67.2 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.0 · 2.0 | Coding | 66.7%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.0 (Terminus-2) | Coding | 66.7%Reported settings & sourceKimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 · 2.1 | Coding | 53.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon | Agentic | 50%Reported settings & sourceKimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| V* (w/ python) | Supporting evidence | 96.9%Reported settings & sourceKimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Vals.ai Financial Agent 1.1 with web search · 1.1 | Supporting evidence | 58.8%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 58.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Vals.ai Financial Agent 1.1 without web search · 1.1 | Supporting evidence | 54%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 54.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| VIBE-V2 · source release snapshot; version not specified | Supporting evidence | 46%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| WideSearch (item-f1) | Supporting evidence | 80.8%Reported settings & sourceKimi K2.6 card, WideSearch (item-f1). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, WideSearch (item-f1), column 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WMT24++ (en→xx) · source release snapshot; version not specified | Supporting evidence | 84.5 score (source scale)Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 84.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |