Compare published benchmarks
Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.
These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.
Gemini 3.1 Pro and GPT-5.4
Weighted benchmark win share: Gemini 3.1 Pro 0.0% · GPT-5.4 100.0%. Based on 7 shared measures out of 32 reported measures for this pair. Exact reported ties contribute half to each model.
Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.
| Benchmark | Gemini 3.1 Pro | GPT-5.4 | Comparison |
|---|---|---|---|
| HLE-Full (w/ tools) | 51.4% Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 5 · reviewed 2026-09-12 | 52.1% Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| BrowseComp | 85.9% Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 5 · reviewed 2026-09-12 | 82.7% Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| BrowseComp (Agent Swarm) | 85.9% Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 5 · reviewed 2026-09-12 | 82.7% Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| DeepSearchQA (f1-score) | 81.9% Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 5 · reviewed 2026-09-12 | 78.6% Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| DeepSearchQA (accuracy) | 60.2% Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 5 · reviewed 2026-09-12 | 63.7% Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| Toolathlon | 48.8% Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 5 · reviewed 2026-09-12 | 54.6% Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| MCPMark | 55.9% Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 5 · reviewed 2026-09-12 | 62.5% Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 3 · reviewed 2026-09-12 | GPT-5.4 · 50.0% weight |
| Claw Eval (pass^3) | 57.8% Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 5 · reviewed 2026-09-12 | 60.3% Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| Claw Eval (pass@3) | 82.9% Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 5 · reviewed 2026-09-12 | 78.4% Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| APEX-Agents | 32% Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 5 · reviewed 2026-09-12 | 33.3% Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| OSWorld-Verified | No result | 75% Kimi K2.6 card, OSWorld-Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, OSWorld-Verified, column 3 · reviewed 2026-09-12 | Missing result |
| Terminal-Bench 2.0 (Terminus-2) | 68.5% Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 5 · reviewed 2026-09-12 | 65.4% Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| SWE-Bench Pro | 54.2% Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 5 · reviewed 2026-09-12 | 57.7% Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| SWE-Bench Multilingual | 76.9% Kimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 5 · reviewed 2026-09-12 | No result | Missing result |
| SWE-Bench Verified | 80.6% Kimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 5 · reviewed 2026-09-12 | No result | Missing result |
| SciCode | 58.9% Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SciCode, column 5 · reviewed 2026-09-12 | 56.6% Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, SciCode, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| OJBench (python) | 70.7% Kimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 5 · reviewed 2026-09-12 | No result | Missing result |
| LiveCodeBench (v6) | 91.7% Kimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 5 · reviewed 2026-09-12 | No result | Missing result |
| HLE-Full | 44.4% Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 5 · reviewed 2026-09-12 | 39.8% Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| AIME 2026 | 98.3% Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 5 · reviewed 2026-09-12 | 99.2% Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| HMMT 2026 (Feb) | 94.7% Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 5 · reviewed 2026-09-12 | 97.7% Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| IMO-AnswerBench | 91% Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 5 · reviewed 2026-09-12 | 91.4% Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| GPQA-Diamond | 94.3% Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 5 · reviewed 2026-09-12 | 92.8% Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| MMMU-Pro | 83% Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 5 · reviewed 2026-09-12 | 81.2% Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| MMMU-Pro (w/ python) | 85.3% Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 5 · reviewed 2026-09-12 | 82.1% Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| CharXiv (RQ) | 80.2% Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 5 · reviewed 2026-09-12 | 82.8% Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
| CharXiv (RQ) (w/ python) | 89.9% Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 5 · reviewed 2026-09-12 | 90% Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
| MathVision | 89.8% Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision, column 5 · reviewed 2026-09-12 | 92% Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision, column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
| MathVision (w/ python) | 95.7% Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 5 · reviewed 2026-09-12 | 96.1% Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
| BabyVision | 51.6% Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 5 · reviewed 2026-09-12 | 49.7% Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation. Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run. Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 3 · reviewed 2026-09-12 | Cited comparator requires original evaluation provenance; not a new matched run. |
| BabyVision (w/ python) | 68.3% Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 5 · reviewed 2026-09-12 | 80.2% Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
| V* (w/ python) | 96.9% Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 5 · reviewed 2026-09-12 | 98.4% Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps. Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission. Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 3 · reviewed 2026-09-12 | GPT-5.4 · 8.3% weight |
Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations