Compare published benchmarks

Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.

These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.

Category weights

Weights are relative. Categories share the total; each benchmark family shares its category equally. Related versions and settings divide their family’s allocation.

Gemini 3.1 Pro and Kimi K3

No unambiguous shared results under these settings. Inspect the available results below or choose another source.

Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.

BenchmarkGemini 3.1 ProKimi K3Comparison
HLE-Full (w/ tools)
Agentic · Higher is better

51.4%

Reported settings & source

Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 5 · reviewed 2026-09-12

No resultMissing result
BrowseComp
Agentic · Higher is better

85.9%

Reported settings & source

Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 5 · reviewed 2026-09-12

No resultMissing result
BrowseComp (Agent Swarm)
Agentic · Higher is better

85.9%

Reported settings & source

Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 5 · reviewed 2026-09-12

No resultMissing result
DeepSearchQA (f1-score)
Agentic · Higher is better

81.9%

Reported settings & source

Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 5 · reviewed 2026-09-12

No resultMissing result
DeepSearchQA (accuracy)
Agentic · Higher is better

60.2%

Reported settings & source

Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 5 · reviewed 2026-09-12

No resultMissing result
Toolathlon
Agentic · Higher is better

48.8%

Reported settings & source

Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 5 · reviewed 2026-09-12

No resultMissing result
MCPMark
Agentic · Higher is better

55.9%

Reported settings & source

Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 5 · reviewed 2026-09-12

No resultMissing result
Claw Eval (pass^3)
Agentic · Higher is better

57.8%

Reported settings & source

Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 5 · reviewed 2026-09-12

No resultMissing result
Claw Eval (pass@3)
Agentic · Higher is better

82.9%

Reported settings & source

Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 5 · reviewed 2026-09-12

No resultMissing result
APEX-Agents
Agentic · Higher is better

32%

Reported settings & source

Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 5 · reviewed 2026-09-12

No resultMissing result
Terminal-Bench 2.0 (Terminus-2)
Coding · Higher is better

68.5%

Reported settings & source

Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 5 · reviewed 2026-09-12

No resultMissing result
SWE-Bench Pro
Coding · Higher is better

54.2%

Reported settings & source

Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 5 · reviewed 2026-09-12

No resultMissing result
SWE-Bench Multilingual
Coding · Higher is better

76.9%

Reported settings & source

Kimi K2.6 card, SWE-Bench Multilingual. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Multilingual, column 5 · reviewed 2026-09-12

No resultMissing result
SWE-Bench Verified
Coding · Higher is better

80.6%

Reported settings & source

Kimi K2.6 card, SWE-Bench Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Verified, column 5 · reviewed 2026-09-12

No resultMissing result
SciCode
Coding · Higher is better

58.9%

Reported settings & source

Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, SciCode, column 5 · reviewed 2026-09-12

No resultMissing result
OJBench (python)
Coding · Higher is better

70.7%

Reported settings & source

Kimi K2.6 card, OJBench (python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, OJBench (python), column 5 · reviewed 2026-09-12

No resultMissing result
LiveCodeBench (v6)
Coding · Higher is better

91.7%

Reported settings & source

Kimi K2.6 card, LiveCodeBench (v6). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, LiveCodeBench (v6), column 5 · reviewed 2026-09-12

No resultMissing result
HLE-Full
Hard reasoning · Higher is better

44.4%

Reported settings & source

Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 5 · reviewed 2026-09-12

No resultMissing result
AIME 2026
Hard reasoning · Higher is better

98.3%

Reported settings & source

Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 5 · reviewed 2026-09-12

No resultMissing result
HMMT 2026 (Feb)
Hard reasoning · Higher is better

94.7%

Reported settings & source

Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 5 · reviewed 2026-09-12

No resultMissing result
IMO-AnswerBench
Hard reasoning · Higher is better

91%

Reported settings & source

Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 5 · reviewed 2026-09-12

No resultMissing result
GPQA-Diamond
Hard reasoning · Higher is better

94.3%

Reported settings & source

Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 5 · reviewed 2026-09-12

No resultMissing result
MMMU-Pro
Multimodal · Higher is better

83%

Reported settings & source

Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 5 · reviewed 2026-09-12

No resultMissing result
MMMU-Pro (w/ python)
Multimodal · Higher is better

85.3%

Reported settings & source

Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 5 · reviewed 2026-09-12

No resultMissing result
CharXiv (RQ)
Multimodal · Higher is better

80.2%

Reported settings & source

Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 5 · reviewed 2026-09-12

No resultMissing result
CharXiv (RQ) (w/ python)
Multimodal · Higher is better

89.9%

Reported settings & source

Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 5 · reviewed 2026-09-12

No resultMissing result
MathVision
Multimodal · Higher is better

89.8%

Reported settings & source

Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision, column 5 · reviewed 2026-09-12

No resultMissing result
MathVision (w/ python)
Multimodal · Higher is better

95.7%

Reported settings & source

Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 5 · reviewed 2026-09-12

No resultMissing result
BabyVision
Multimodal · Higher is better

51.6%

Reported settings & source

Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 5 · reviewed 2026-09-12

No resultMissing result
BabyVision (w/ python)
Multimodal · Higher is better

68.3%

Reported settings & source

Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 5 · reviewed 2026-09-12

No resultMissing result
V* (w/ python)
Multimodal · Higher is better

96.9%

Reported settings & source

Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 5 · reviewed 2026-09-12

No resultMissing result

Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations