RankingGPT-5.4

GPT-5.4

Data updated 12 Sept 2026

32 published benchmark measures · 8 benchmark families contribute across 5 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-5.4 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-5.4 · Hard reasoning: 65.5 · SupportedGPT-5.4 · Coding: 8.0 · SupportedGPT-5.4 · Human pref: 36.1 · SupportedGPT-5.4 · Multimodal: 52.0 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoning2 families · 1 with independent evidence · Supported65.5

    2 core families; 51 direct opponents across 11 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 62.6–71.1; 2/7 scenarios unsupported. Without one publisher: 57.7–80.5; 2/16 unsupported. Smoothing check: 62.9–66.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • ARC-AGI: 76.6 observed win share
      arcprize.org · Source 1
    • GPQA: 100.0 observed win share
      Microsoft · Source 1
    Coding3 families · 2 with independent evidence · Supported8.0

    3 core families; 31 direct opponents across 10 labs.

    Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 6.8–25.3; 2/7 scenarios unsupported. Without one publisher: 2.2–30.0; 1/20 unsupported. Smoothing check: 5.5–14.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • SWE-bench Pro: 58.3 observed win share
      Microsoft, scale.com · Source 1 Source 2
    • DeepSWE: 22.2 observed win share
      deepswe.datacurve.ai · Source 1
    • Terminal-Bench: 0.0 observed win share
      Moonshot AI · Source 1
    Human pref1 families · 1 with independent evidence · Supported36.1

    1 core families; 114 direct opponents across 19 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 28.3–42.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Human preference (Arena): 85.1 observed win share
      lmarena.ai · Source 1
    Knowledge1 families · 1 with independent evidence · PreliminaryUnknown

    1 core families; 51 direct opponents across 12 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • MMLU-Pro: 94.1 observed win share
      huggingface.co · Source 1
    Multimodal1 families · 0 with independent evidence · Preliminary52.0

    1 core families; 2 direct opponents across 2 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • CharXiv: 100.0 observed win share
      Moonshot AI · Source 1
    Long contextNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 5 effort levels across 51 benchmark/harness combinations →

      Reported effort · XHigh + unspecified

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      8 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      OpenAI
      Catalog status
      active
      Availability
      Public provider catalog; account and region restrictions may apply
      Family
      gpt-5.4
      Released
      Context
      License
      proprietary
      Model card
      https://developers.openai.com/api/docs/models/all
      Default Capability family coverage
      Documented-evidence family coverage/badge/gpt-5.4.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      AIME 2026Hard reasoning99.2%
      Reported settings & source

      Kimi K2.6 card, AIME 2026. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, AIME 2026, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      APEX-AgentsAgentic33.3%
      Reported settings & source

      Kimi K2.6 card, APEX-Agents. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, APEX-Agents, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      ARC-AGI-2Hard reasoning74%ARC Prize verified

      Effort: Not specified

      contributes to capability
      official board2026-03-04
      BabyVisionSupporting evidence49.7%
      Reported settings & source

      Kimi K2.6 card, BabyVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, BabyVision, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      BabyVision (w/ python)Supporting evidence80.2%
      Reported settings & source

      Kimi K2.6 card, BabyVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, BabyVision (w/ python), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      BrowseCompAgentic82.7%
      Reported settings & source

      Kimi K2.6 card, BrowseComp. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, BrowseComp, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      BrowseComp (Agent Swarm)Agentic82.7%
      Reported settings & source

      Kimi K2.6 card, BrowseComp (Agent Swarm). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, BrowseComp (Agent Swarm), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      CharXiv (RQ)Multimodal82.8%
      Reported settings & source

      Kimi K2.6 card, CharXiv (RQ). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      contributes to capability
      lab self-reportReviewed 2026-09-12
      CharXiv (RQ) (w/ python)Supporting evidence90%
      Reported settings & source

      Kimi K2.6 card, CharXiv (RQ) (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, CharXiv (RQ) (w/ python), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Claw Eval (pass@3)Supporting evidence78.4%
      Reported settings & source

      Kimi K2.6 card, Claw Eval (pass@3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass@3), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Claw Eval (pass^3)Supporting evidence60.3%
      Reported settings & source

      Kimi K2.6 card, Claw Eval (pass^3). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, Claw Eval (pass^3), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DeepSearchQA (accuracy)Supporting evidence63.7%
      Reported settings & source

      Kimi K2.6 card, DeepSearchQA (accuracy). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (accuracy), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DeepSearchQA (f1-score)Supporting evidence78.6%
      Reported settings & source

      Kimi K2.6 card, DeepSearchQA (f1-score). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, DeepSearchQA (f1-score), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DeepSWE v1.1Coding51.8%DeepSWE v1.1 reported

      Effort: Not specified

      contributes to capability
      official board2026-09-03
      GPQA Diamond · source release snapshot; version not specifiedHard reasoning92.8%
      Reported settings & source

      MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

      First-party reported result; comparator results retain the source evaluation setup.

      MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GPQA-DiamondSupporting evidence92.8%
      Reported settings & source

      Kimi K2.6 card, GPQA-Diamond. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, GPQA-Diamond, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HLE-FullHard reasoning39.8%
      Reported settings & source

      Kimi K2.6 card, HLE-Full. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, HLE-Full, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      HLE-Full (w/ tools)Hard reasoning52.1%
      Reported settings & source

      Kimi K2.6 card, HLE-Full (w/ tools). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, HLE-Full (w/ tools), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      HMMT 2026 (Feb)Supporting evidence97.7%
      Reported settings & source

      Kimi K2.6 card, HMMT 2026 (Feb). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, HMMT 2026 (Feb), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      IMO-AnswerBenchSupporting evidence91.4%
      Reported settings & source

      Kimi K2.6 card, IMO-AnswerBench. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Reasoning: 98304 generation tokens; HLE full set.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, IMO-AnswerBench, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      LMArena Text ArenaHuman pref1466LMArena Text

      Effort: Not specified

      contributes to capability
      official board2026-09-11
      MathVisionSupporting evidence92%
      Reported settings & source

      Kimi K2.6 card, MathVision. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, MathVision, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MathVision (w/ python)Supporting evidence96.1%
      Reported settings & source

      Kimi K2.6 card, MathVision (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, MathVision (w/ python), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MCPMarkSupporting evidence62.5%
      Reported settings & source

      Kimi K2.6 card, MCPMark. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, MCPMark, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MMLU-ProKnowledge87.5%MMLU-Pro reported

      Effort: Not specified

      contributes to capability
      official board2026-09-11
      MMMU-ProMultimodal81.2%
      Reported settings & source

      Kimi K2.6 card, MMMU-Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      MMMU-Pro (w/ python)Supporting evidence82.1%
      Reported settings & source

      Kimi K2.6 card, MMMU-Pro (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, MMMU-Pro (w/ python), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      OSWorld-VerifiedAgentic75%
      Reported settings & source

      Kimi K2.6 card, OSWorld-Verified. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, OSWorld-Verified, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      SciCodeCoding56.6%
      Reported settings & source

      Kimi K2.6 card, SciCode. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, SciCode, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      SWE-bench ProCoding59.1%SWE-bench Pro reported

      Effort: Not specified

      contributes to capability
      official board2026-04-08
      SWE-Bench ProCoding57.7%
      Reported settings & source

      Kimi K2.6 card, SWE-Bench Pro. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, SWE-Bench Pro, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      SWE-bench Pro · source release snapshot; version not specifiedCoding57.7%
      Reported settings & source

      MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

      First-party reported result; comparator results retain the source evaluation setup.

      MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench 2.0 · 2.0Coding75.1%
      Reported settings & source

      MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

      First-party reported result; comparator results retain the source evaluation setup.

      MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Microsoft report Table 11 and section 4.1: MAI Terminal-Bench removes timeouts and uses a minimal ReAct harness; comparator values are cited from other model releases. A common evaluation protocol is not established.

      Reviewed 2026-09-06
      Terminal-Bench 2.0 (Terminus-2)Coding65.4%
      Reported settings & source

      Kimi K2.6 card, Terminal-Bench 2.0 (Terminus-2). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Coding: average of ten runs; Terminal uses Terminus2 preserve-thinking; SWE uses in-house SWE-agent.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, Terminal-Bench 2.0 (Terminus-2), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      contributes to capability
      lab self-reportReviewed 2026-09-12
      ToolathlonAgentic54.6%
      Reported settings & source

      Kimi K2.6 card, Toolathlon. Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Cited from another official report; does not establish a new same-protocol evaluation.

      Comparison limit: Cited comparator requires original evaluation provenance; not a new matched run.

      Kimi K2.6 official model card · Evaluation Results table, Toolathlon, column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Cited comparator requires original evaluation provenance; not a new matched run.

      Reviewed 2026-09-12
      V* (w/ python)Supporting evidence98.4%
      Reported settings & source

      Kimi K2.6 card, V* (w/ python). Thinking mode; temperature 1.0, top-p 1.0, context 262144. Comparator settings: GPT-5.4 xhigh, Opus4.6 max, Gemini3.1Pro high. See benchmark-specific footnotes. Only own K2.6 and starred re-evaluations can form new comparisons. Vision: 98304 tokens, average of three runs; Python condition uses 65536 tokens per step, at most50steps.

      Source SHA256 95db3be1d0473e482c7aa901f237ad7af8bb7929625a777b27428081bb8ea9f7. Own measurement or explicitly starred same-condition re-evaluation. Assess benchmark-specific protocol before admission.

      Kimi K2.6 official model card · Evaluation Results table, V* (w/ python), column 3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      GDPval-AAAgentic
      GPQA DiamondHard reasoning
      Humanity's Last ExamHard reasoning
      LiveCodeBenchCoding
      OSWorld-VerifiedAgentic
      SWE-bench VerifiedAgentic
      Terminal-Bench 2.1Agentic