RankingClaude Opus 4.8

Claude Opus 4.8

Data updated 12 Sept 2026

108 published benchmark measures · 22 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Opus 4.8 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Claude Opus 4.8 · Agentic: 59.6 · SupportedClaude Opus 4.8 · Hard reasoning: 38.7 · SupportedClaude Opus 4.8 · Coding: 57.0 · SupportedClaude Opus 4.8 · Human pref: 44.5 · SupportedClaude Opus 4.8 · Multimodal: 58.4 · SupportedClaude Opus 4.8 · Long context: 63.1 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic9 families · 0 with independent evidence · Supported59.6

9 core families; 15 direct opponents across 8 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 56.2–62.8; 0/11 scenarios unsupported. Without one publisher: 58.3–61.2; 0/15 unsupported. Smoothing check: 57.7–59.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 1 with independent evidence · Supported38.7

4 core families; 52 direct opponents across 10 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 31.9–43.8; 0/7 scenarios unsupported. Without one publisher: 36.4–40.0; 0/16 unsupported. Smoothing check: 37.8–41.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding3 families · 1 with independent evidence · Supported57.0

3 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 46.0–66.9; 0/7 scenarios unsupported. Without one publisher: 51.2–59.1; 0/20 unsupported. Smoothing check: 55.9–57.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported44.5

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 40.9–47.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 90.4 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal2 families · 0 with independent evidence · Supported58.4

    2 core families; 6 direct opponents across 5 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 57.9–59.2; 2/8 scenarios unsupported. Without one publisher: 47.1–59.2; 1/8 unsupported. Smoothing check: 56.7–59.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context3 families · 0 with independent evidence · Supported63.1

    3 core families; 6 direct opponents across 2 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 53.9–65.9; 1/5 scenarios unsupported. Without one publisher: 63.1–63.1; 3/5 unsupported. Smoothing check: 62.5–63.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 6 effort levels across 73 benchmark/harness combinations →

    Reported effort · Mixed settings

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    22 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Anthropic
    Catalog status
    active
    Availability
    Public provider catalog; account and region restrictions may apply
    Family
    Claude Opus
    Released
    Context
    License
    proprietary
    Model card
    https://platform.claude.com/docs/en/about-claude/model-deprecations
    Default Capability family coverage
    Documented-evidence family coverage/badge/claude-opus-4-8.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    $OneMillion-Bench (expert score) · source release snapshot; version not specifiedSupporting evidence41.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1354
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-LCR · source release snapshot; version not specifiedLong context67.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Agents' Last Exam · not specifiedSupporting evidence45.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence27%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence25.7%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence25.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence27%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence45.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AndroidBench · source release snapshot; version not specifiedSupporting evidence69.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic39.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    ARC-AGI-2Hard reasoning72.1%ARC Prize verified

    Effort: Not specified

    contributes to capability
    official board2026-06-01
    ARC-AGI-3 · 3Hard reasoning1.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; high reasoning, not max

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: High

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence72.5 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence55.7 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Automation-Bench (Pass@1) · source release snapshot; version not specifiedAgentic27.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench · not specifiedAgentic15.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · source release snapshot; version not specifiedAgentic27.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench (Public) · source release snapshot; version not specifiedAgentic27.2%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench (v1.0.6) · v1.0.6Agentic41%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    BabyVision w/ python · source release snapshot; version not specifiedSupporting evidence81.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BabyVision with tools · source release snapshot; version not specifiedSupporting evidence81.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD · not specifiedSupporting evidence27.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD (python tool) · not specifiedSupporting evidence51.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Big Finance Bench · not specifiedSupporting evidence44%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseComp · not specifiedAgentic84.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic84.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal80.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal89.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv Reasoning with tools · source release snapshot; version not specifiedMultimodal89.9%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    CorpFin v2 · source release snapshot; version not specifiedSupporting evidence66.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CoWorkBench · source release snapshot; version not specifiedSupporting evidence72.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CritPt · source release snapshot; version not specifiedSupporting evidence20.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Cybergym · source release snapshot; version not specifiedSupporting evidence78.3%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CyberGym · source release snapshot; version not specifiedSupporting evidence78.1%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DeepSearchQA · source release snapshot; version not specifiedAgentic84.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSearchQA (F1) · source release snapshot; version not specifiedAgentic93.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSearchQA (F1) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE · source release snapshot; version not specifiedCoding59%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE · source release snapshot; version not specifiedCoding58%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE (v1.1) · v1.1Coding58%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE 1.1 · 1.1Coding59%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE v1.1Coding59%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding59%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding59%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DSBench-FullStack † · source release snapshot; version not specifiedSupporting evidence71.6%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result. † source footnote applies.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DSBench-Hard † · source release snapshot; version not specifiedSupporting evidence71.7%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result. † source footnote applies.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitBench · not specifiedSupporting evidence40%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitBench · source release snapshot; version not specifiedSupporting evidence40%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence80 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence120 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence53.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence53.9%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierMath Tier 1-3 (v2) · v2Hard reasoning80%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning56.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence70%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence66.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence66.5%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    Externally evaluated result, attributed in the source footnote.

    Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison.

    zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    gdp.pdf · not specifiedSupporting evidence22.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval-AA v2 · source release snapshot; version not specifiedAgentic1588
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    Externally evaluated result, attributed in the source footnote.

    Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

    zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 · v2Agentic1600
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1593
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 Elo · source release snapshot; version not specifiedAgentic1600
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GeneBench Pro · not specifiedSupporting evidence16%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GPQA Diamond · not specifiedHard reasoning92%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning92%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning91%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GraphWalks BFS 1mil f1 · not specifiedLong context68.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GraphWalks BFS 256k f1 · not specifiedLong context85.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Harvey Lab-AA · source release snapshot; version not specifiedSupporting evidence91.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench · source release snapshot; version not specifiedSupporting evidence52.4%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench Professional · not specifiedSupporting evidence53%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · source release snapshot; version not specifiedSupporting evidence55.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE · source release snapshot; version not specifiedHard reasoning45.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning49.8%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning57.9%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE w/ tools · source release snapshot; version not specifiedHard reasoning57.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE w/ Tools · source release snapshot; version not specifiedHard reasoning57.9%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE with tools · source release snapshot; version not specifiedHard reasoning57.9%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE without tools · source release snapshot; version not specifiedHard reasoning49.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE-Full · source release snapshot; version not specifiedHard reasoning49.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning57.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    IFBench · source release snapshot; version not specifiedSupporting evidence62.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence48.4%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence48.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence48.4%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Kimi Code Bench 2.0 · 2.0Supporting evidence71.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Legal Research Bench · source release snapshot; version not specifiedSupporting evidence43.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    LifeSciBench · not specifiedSupporting evidence53.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LMArena Text ArenaHuman pref1473LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    LongBench v2 · source release snapshot; version not specifiedLong context69.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Management Consulting Tasks (Internal) · not specifiedSupporting evidence31.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MathVision · source release snapshot; version not specifiedSupporting evidence86.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MathVision · source release snapshot; version not specifiedSupporting evidence97.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MCP-Atlas · source release snapshot; version not specifiedAgentic83.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MCPAtlas · source release snapshot; version not specifiedAgentic82.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MCPMark-Verified · source release snapshot; version not specifiedSupporting evidence76.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence42.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence42.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal78.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal82.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMVU · source release snapshot; version not specifiedSupporting evidence79.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MRCR v2 256K (8-needle) · source release snapshot; version not specifiedLong context83.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    NL2Repo · source release snapshot; version not specifiedSupporting evidence69.7%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: NL2Repo · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    NL2Repo · source release snapshot; version not specifiedSupporting evidence69.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    NL2Repo-Bench · source release snapshot; version not specifiedSupporting evidence69.4%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: NL2Repo-Bench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OfficeQA Pro · source release snapshot; version not specifiedAgentic63.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OmniDocBench · source release snapshot; version not specifiedSupporting evidence87.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OSWorld 2.0 · 2.0Agentic54.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 · 2.0Agentic55.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld 2.0 binary without exec · 2.0Agentic20.6%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 partial without exec · 2.0Agentic54.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld Verified · source release snapshot; version not specifiedAgentic83.4%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld-Verified · source release snapshot; version not specifiedAgentic83.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    PaperBench · source release snapshot; version not specifiedSupporting evidence80.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PerceptionBench · source release snapshot; version not specifiedSupporting evidence47.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PLawBench · source release snapshot; version not specifiedSupporting evidence69.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence34.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official PostTrainBench results; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence32.9%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PRBench-Finance · source release snapshot; version not specifiedSupporting evidence51.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PRBench-Legal · source release snapshot; version not specifiedSupporting evidence52.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench · source release snapshot; version not specifiedSupporting evidence71.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench (Almost Solved) · source release snapshot; version not specifiedSupporting evidence15.5%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenQoderBench · source release snapshot; version not specifiedSupporting evidence62.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenReactBench · source release snapshot; version not specifiedSupporting evidence1694
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSVGBench · source release snapshot; version not specifiedSupporting evidence1648
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSWEBench · source release snapshot; version not specifiedSupporting evidence84%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ResearchRubrics · source release snapshot; version not specifiedSupporting evidence73.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SaaS-Bench · source release snapshot; version not specifiedSupporting evidence56.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SciCode · source release snapshot; version not specifiedCoding53.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    SkillsBench · source release snapshot; version not specifiedSupporting evidence65.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SpreadsheetBench 2 · source release snapshot; version not specifiedSupporting evidence31.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-Bench Pro · not specifiedCoding69.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding69.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-bench Pro · source release snapshot; version not specifiedCoding69.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-Marathon · source release snapshot; version not specifiedSupporting evidence40%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-Marathon (v1.1) · v1.1Supporting evidence48.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding84.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding85%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding85%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal Bench 3.0 · 3.0Coding21.1%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal-Bench 2.1 · 2.1Coding78.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding84.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Terminal-Bench 2.1 · 2.1Coding82.7%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon · not specifiedAgentic59.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / Claude Opus 4.8 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified · source release snapshot; version not specifiedAgentic76.2%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon Verified · source release snapshot; version not specifiedAgentic76.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified (Pass@1) · source release snapshot; version not specifiedAgentic76.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic76.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic76.2%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Video-MME (w. sub) · source release snapshot; version not specifiedSupporting evidence86%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WebArena Verified · source release snapshot; version not specifiedAgentic71.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    WideSearch · source release snapshot; version not specifiedSupporting evidence72.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Item-F1 over4runs; Qwen-Agent for Qwen,ClaudeCode for comparators.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: WideSearch · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WorkSpaceBench · source release snapshot; version not specifiedSupporting evidence66.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WorldVQA ForceAnswer · source release snapshot; version not specifiedSupporting evidence39.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence17%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence34%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    τ³-Banking · source release snapshot; version not specifiedSupporting evidence27.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    GDPval-AAAgentic
    GPQA DiamondHard reasoning
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    OSWorld-VerifiedAgentic
    SWE-bench ProAgentic
    SWE-bench VerifiedAgentic
    Terminal-Bench 2.1Agentic