RankingKimi K3

Kimi K3

Data updated 12 Sept 2026

67 published benchmark measures · 15 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Kimi K3 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Kimi K3 · Agentic: 86.0 · SupportedKimi K3 · Hard reasoning: 52.6 · SupportedKimi K3 · Coding: 72.1 · SupportedKimi K3 · Human pref: 52.7 · SupportedKimi K3 · Multimodal: 76.4 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic7 families · 1 with independent evidence · Supported86.0

7 core families; 46 direct opponents across 18 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 75.8–88.1; 0/11 scenarios unsupported. Without one publisher: 79.4–89.8; 0/15 unsupported. Smoothing check: 77.7–89.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning3 families · 2 with independent evidence · Supported52.6

3 core families; 59 direct opponents across 11 labs.

Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 41.5–70.3; 0/7 scenarios unsupported. Without one publisher: 42.6–56.6; 0/16 unsupported. Smoothing check: 51.6–53.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 1 with independent evidence · Supported72.1

2 core families; 27 direct opponents across 9 labs.

Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 66.7–72.9; 2/7 scenarios unsupported. Without one publisher: 64.6–76.4; 0/20 unsupported. Smoothing check: 67.3–74.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported52.7

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 51.3–54.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 94.7 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal2 families · 0 with independent evidence · Supported76.4

    2 core families; 4 direct opponents across 2 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 75.4–77.8; 2/8 scenarios unsupported. Without one publisher: 75.1–79.1; 1/8 unsupported. Smoothing check: 70.8–79.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long contextNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 4 effort levels across 41 benchmark/harness combinations →

      Reported effort · Max + unspecified

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      15 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      Moonshot
      Catalog status
      api-listed
      Availability
      Documented provider API and open weights for self-hosting
      Family
      Kimi
      Released
      Context
      1,000,000 tokens
      License
      Model card
      https://huggingface.co/moonshotai/Kimi-K3
      Default Capability family coverage
      Documented-evidence family coverage/badge/kimi-k3.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1548
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-LCR · source release snapshot; version not specifiedLong context74.7%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      Reviewed 2026-09-12
      Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence28.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence27.6%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence27.6%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      APEX-Agents · source release snapshot; version not specifiedAgentic41%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

      Reviewed 2026-09-12
      ARC-AGI-2Hard reasoning60.4%ARC Prize verified

      Effort: Not specified

      contributes to capability
      official board2026-07-16
      Artificial Analysis Intelligence IndexSupporting evidence60 AA Intelligence Index

      Effort: Not specified

      official board

      Supporting evidence outside the reviewed capability core

      2026-08-12
      AutomationBench · source release snapshot; version not specifiedAgentic30.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      AutomationBench (Public) · source release snapshot; version not specifiedAgentic30.8%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      AutomationBench (v1.0.6) · v1.0.6Agentic46.7%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      BabyVision w/ python · source release snapshot; version not specifiedSupporting evidence85.7%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      BrowseCompAgentic91.2%BrowseComp reported

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-07-16
      BrowseComp · source release snapshot; version not specifiedAgentic91.2%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-12
      CharXiv (RQ) · source release snapshot; version not specifiedMultimodal84.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      CharXiv (RQ) · source release snapshot; version not specifiedMultimodal91.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      CorpFin v2 · source release snapshot; version not specifiedSupporting evidence71.6%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      CritPt · source release snapshot; version not specifiedSupporting evidence23.4%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Cybergym · source release snapshot; version not specifiedSupporting evidence80%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      CyberGym · source release snapshot; version not specifiedSupporting evidence80%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DeepSearchQA (F1) · source release snapshot; version not specifiedAgentic95%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: DeepSearchQA (F1) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      DeepSWE · source release snapshot; version not specifiedCoding67.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      DeepSWE · source release snapshot; version not specifiedCoding67.5%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      DeepSWE (v1.1) · v1.1Coding67.5%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      DeepSWE v1.1Coding68.5%DeepSWE v1.1 reported

      Effort: Not specified

      contributes to capability
      official board2026-09-03
      DSBench-FullStack † · source release snapshot; version not specifiedSupporting evidence73.7%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result. † source footnote applies.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DSBench-Hard † · source release snapshot; version not specifiedSupporting evidence63%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result. † source footnote applies.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ExploitBench · source release snapshot; version not specifiedSupporting evidence32.2%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence36 count
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence70 count
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence54.4%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      FrontierSWE · source release snapshot; version not specifiedSupporting evidence81.2%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      GDPval-AAAgentic1569Artificial Analysis GDPval-AA

      Effort: Not specified

      contributes to capability
      official board2026-09-12
      GDPval-AA v2 · source release snapshot; version not specifiedAgentic1682
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      Externally evaluated result, attributed in the source footnote.

      Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

      zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

      Reviewed 2026-09-12
      GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1686
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      Reviewed 2026-09-12
      GPQA DiamondHard reasoning93.535%GPQA Diamond reported

      Effort: Not specified

      contributes to capability
      independent repro2026-09-11
      GPQA Diamond · source release snapshot; version not specifiedHard reasoning93.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Harvey Lab-AA · source release snapshot; version not specifiedSupporting evidence94.6%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning43.5%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning56%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      HLE w/ Tools · source release snapshot; version not specifiedHard reasoning59.8%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      HLE-Full · source release snapshot; version not specifiedHard reasoning43.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      HLE-Full · source release snapshot; version not specifiedHard reasoning56%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Humanity's Last ExamHard reasoning43.5%HLE no tools

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-07-16
      JobBench · source release snapshot; version not specifiedSupporting evidence54.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Kimi Code Bench 2.0 · 2.0Supporting evidence72.9%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Legal Research Bench · source release snapshot; version not specifiedSupporting evidence44.2%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      LMArena Text ArenaHuman pref1485LMArena Text

      Effort: Not specified

      contributes to capability
      official board2026-09-11
      MathVision · source release snapshot; version not specifiedSupporting evidence94.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MathVision · source release snapshot; version not specifiedSupporting evidence97.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MCP-Atlas · source release snapshot; version not specifiedAgentic84.2%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      MCPMark-Verified · source release snapshot; version not specifiedSupporting evidence94.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence48.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MMMU-Pro · source release snapshot; version not specifiedMultimodal81.6%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      MMMU-Pro · source release snapshot; version not specifiedMultimodal83.4%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      MMVU · source release snapshot; version not specifiedSupporting evidence82.1%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      NL2Repo · source release snapshot; version not specifiedSupporting evidence58%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      OfficeQA Pro · source release snapshot; version not specifiedAgentic63.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OmniDocBench · source release snapshot; version not specifiedSupporting evidence91.1%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      OSWorld 2.0 · 2.0Agentic58.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OSWorld-Verified · source release snapshot; version not specifiedAgentic84.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      PerceptionBench · source release snapshot; version not specifiedSupporting evidence58.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      PostTrainBench · source release snapshot; version not specifiedSupporting evidence36.6%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      PostTrainBench · source release snapshot; version not specifiedSupporting evidence32%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ProgramBench · source release snapshot; version not specifiedSupporting evidence77.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ProgramBench (Almost Solved) · source release snapshot; version not specifiedSupporting evidence17.5%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ResearchRubrics · source release snapshot; version not specifiedSupporting evidence76.2%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SaaS-Bench · source release snapshot; version not specifiedSupporting evidence60.1%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SciCode · source release snapshot; version not specifiedCoding58.7%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      Reviewed 2026-09-12
      SpreadsheetBench 2 · source release snapshot; version not specifiedSupporting evidence34.8%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SWE-Marathon · source release snapshot; version not specifiedSupporting evidence42%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SWE-Marathon (v1.1) · v1.1Supporting evidence48.1%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Terminal Bench 2.1 · 2.1Coding88.3%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Terminal Bench 2.1 · 2.1Coding88.3%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Terminal Bench 3.0 · 3.0Coding17.4%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Terminal-Bench 2.1Coding88.3%Terminal-Bench 2.1 reported

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-07-16
      Terminal-Bench 2.1 · 2.1Coding88.3%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-12
      Toolathlon Verified · source release snapshot; version not specifiedAgentic76.5%
      Reported settings & source

      GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

      First-party reported result.

      zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Toolathlon-Verified · source release snapshot; version not specifiedAgentic76.5%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Toolathlon-Verified · source release snapshot; version not specifiedAgentic76.5%
      Reported settings & source

      DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

      First-party reported result.

      deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Video-MME (w. sub) · source release snapshot; version not specifiedSupporting evidence90%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      VulcanBench v3Supporting evidence73.9%VulcanBench v3 bare-bones API · Report 08 · max

      Effort: Not specified

      official board

      Supporting evidence outside the reviewed capability core

      2026-07-19
      WorldVQA ForceAnswer · source release snapshot; version not specifiedSupporting evidence51%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence23%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence41%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      First-party reported result.

      moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      τ³-Banking · source release snapshot; version not specifiedSupporting evidence33.4%
      Reported settings & source

      Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

      Cited result; see benchmark-specific Evaluation Details.

      Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

      moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      LiveCodeBenchCoding
      MMLU-ProKnowledge
      OSWorld-VerifiedAgentic
      SWE-bench ProAgentic
      SWE-bench VerifiedAgentic