RankingGPT-5.5

GPT-5.5

Data updated 12 Sept 2026

110 published benchmark measures · 23 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-5.5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-5.5 · Agentic: 31.4 · SupportedGPT-5.5 · Hard reasoning: 55.0 · SupportedGPT-5.5 · Coding: 45.5 · SupportedGPT-5.5 · Human pref: 47.7 · SupportedGPT-5.5 · Multimodal: 45.5 · SupportedGPT-5.5 · Long context: 38.5 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic10 families · 0 with independent evidence · Supported31.4

10 core families; 16 direct opponents across 7 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 27.4–43.5; 0/11 scenarios unsupported. Without one publisher: 23.5–44.2; 0/15 unsupported. Smoothing check: 28.2–37.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 1 with independent evidence · Supported55.0

4 core families; 48 direct opponents across 10 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 40.8–66.0; 0/7 scenarios unsupported. Without one publisher: 38.2–58.6; 0/16 unsupported. Smoothing check: 53.3–56.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding4 families · 1 with independent evidence · Supported45.5

4 core families; 32 direct opponents across 10 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 38.1–57.0; 0/7 scenarios unsupported. Without one publisher: 43.7–47.8; 0/20 unsupported. Smoothing check: 45.0–46.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported47.7

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 46.0–48.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 92.1 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal2 families · 0 with independent evidence · Supported45.5

    2 core families; 12 direct opponents across 6 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 45.5–45.6; 2/8 scenarios unsupported. Without one publisher: 41.1–55.3; 0/8 unsupported. Smoothing check: 45.5–46.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context2 families · 0 with independent evidence · Supported38.5

    2 core families; 6 direct opponents across 4 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 38.5–39.3; 2/5 scenarios unsupported. Without one publisher: 38.5–40.3; 2/5 unsupported. Smoothing check: 37.0–41.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 6 effort levels across 74 benchmark/harness combinations →

    Reported effort · XHigh + unspecified

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    23 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    OpenAI
    Catalog status
    active
    Availability
    Public provider catalog; account and region restrictions may apply
    Family
    gpt-5.5
    Released
    Context
    License
    proprietary
    Model card
    https://developers.openai.com/api/docs/models/all
    Default Capability family coverage
    Documented-evidence family coverage/badge/gpt-5.5.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1158
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-LCR · source release snapshot; version not specifiedLong context74.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Agents' Last Exam · not specifiedSupporting evidence46.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence26.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic38.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic41.7%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-2Hard reasoning85%ARC Prize verified

    Effort: Not specified

    contributes to capability
    official board2026-04-22
    ARC-AGI-3 · 3Hard reasoning0.43%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence76.4 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence54.8 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AutomationBench · not specifiedAgentic12.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · source release snapshot; version not specifiedAgentic22.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    BabyVision w/ python · source release snapshot; version not specifiedSupporting evidence83.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BabyVision with tools · source release snapshot; version not specifiedSupporting evidence83.6%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BankerToolBench · source release snapshot; version not specifiedSupporting evidence70%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD · not specifiedSupporting evidence44.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD (python tool) · not specifiedSupporting evidence55.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Big Finance Bench · not specifiedSupporting evidence49%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseComp · not specifiedAgentic84.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic84.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    BrowseComp · source release snapshot; version not specifiedAgentic84.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Capture-the-Flag Challenges · not specifiedSupporting evidence88.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / Capture-the-Flag Challenges / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal84.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal89%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv Reasoning with tools · source release snapshot; version not specifiedMultimodal84.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    CL-bench · source release snapshot; version not specifiedSupporting evidence25.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CorpFin v2 · source release snapshot; version not specifiedSupporting evidence68.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CritPt · source release snapshot; version not specifiedSupporting evidence27.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DeepSearchQA · source release snapshot; version not specifiedAgentic87.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE · source release snapshot; version not specifiedCoding67%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE v1.1Coding67%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding67%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding67%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ExploitBench · not specifiedSupporting evidence47.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym · not specifiedSupporting evidence15.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; six-hour evaluation cap; alpha API latency rescaled to public API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitGym / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence51.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence51.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierMath Tier 1-3 (v2) · v2Hard reasoning85.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning72.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence64.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    gdp.pdf · not specifiedSupporting evidence26%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval rubrics · source release snapshot; version not specifiedAgentic80.7%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 · v2Agentic1494
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1491
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 Elo · source release snapshot; version not specifiedAgentic1494
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GeneBench Pro · not specifiedSupporting evidence12%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GPQA Diamond · not specifiedHard reasoning93.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning93.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GraphWalks BFS 1mil f1 · not specifiedLong context45.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GraphWalks BFS 256k f1 · not specifiedLong context73.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Harvey Lab-AA · source release snapshot; version not specifiedSupporting evidence86.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench · not specifiedSupporting evidence56.5 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 2313 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.5 / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench · not specifiedSupporting evidence58.4 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 2313 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.5 / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Consensus · not specifiedSupporting evidence95.6 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 2259 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.5 / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Consensus · not specifiedSupporting evidence95.7 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 2259 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.5 / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Hard · not specifiedSupporting evidence31.5 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 2289 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.5 / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Hard · not specifiedSupporting evidence33.8 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 2289 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.5 / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence49.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence51.8 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 3818 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.5 / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence57.2 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 3818 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.5 / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · source release snapshot; version not specifiedSupporting evidence51.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE with tools · source release snapshot; version not specifiedHard reasoning52.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE without tools · source release snapshot; version not specifiedHard reasoning44.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE-Full · source release snapshot; version not specifiedHard reasoning41.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning52.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    IMO 2025 · source release snapshot; version not specifiedSupporting evidence76.1%
    Reported settings & source

    MiniMax points out of42; comparator percentages as printed

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Research Debugging Evaluation · not specifiedSupporting evidence50%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / Internal Research Debugging Evaluation / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    JobBench · source release snapshot; version not specifiedSupporting evidence38.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence38.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    KernelBench Hard · source release snapshot; version not specifiedSupporting evidence20.9%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    KernelGen 1P · not specifiedSupporting evidence29.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / KernelGen 1P / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Kimi Code Bench 2.0 · 2.0Supporting evidence69%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Legal Research Bench · source release snapshot; version not specifiedSupporting evidence40.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    LifeSciBench · not specifiedSupporting evidence50.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LiveSQLBench · source release snapshot; version not specifiedSupporting evidence40.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LMArena Text ArenaHuman pref1476LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    Management Consulting Tasks (Internal) · not specifiedSupporting evidence31.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MathVision · source release snapshot; version not specifiedSupporting evidence92.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MathVision · source release snapshot; version not specifiedSupporting evidence96.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MCP-Atlas · source release snapshot; version not specifiedAgentic82.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MCPAtlas · source release snapshot; version not specifiedAgentic75.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MCPAtlas · source release snapshot; version not specifiedAgentic75.3%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MCPMark-Verified · source release snapshot; version not specifiedSupporting evidence92.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MedChemBench (Internal) · not specifiedSupporting evidence35.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / MedChemBench (Internal) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence35.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MMMU Pro (no tools) · not specifiedMultimodal81.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; no tools

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMMU Pro (with tools) · not specifiedMultimodal83.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; tools enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (with tools) / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMMU-Pro · source release snapshot; version not specifiedMultimodal81.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal83.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal81.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMVU · source release snapshot; version not specifiedSupporting evidence81.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MRCR v2 1M 8-needle · source release snapshot; version not specifiedLong context74%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    NanoGPT · not specifiedSupporting evidence2.65%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / NanoGPT / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    NL2Repo · source release snapshot; version not specifiedSupporting evidence52.9%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OfficeQA Pro · source release snapshot; version not specifiedAgentic60.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OfficeQA Pro · source release snapshot; version not specifiedAgentic52.6%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OmniDocBench · source release snapshot; version not specifiedSupporting evidence89.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OmniDocBench · source release snapshot; version not specifiedSupporting evidence87.5%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OpenAI MRCR v2 8-needle 256K-512K · v2Long context81.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 256K-512K

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OpenAI MRCR v2 8-needle 512K-1M · v2Long context74%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 512K-1M

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 · 2.0Agentic47.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 · 2.0Agentic49.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld 2.0 binary without exec · 2.0Agentic13.9%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 partial without exec · 2.0Agentic47.5%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld Verified · source release snapshot; version not specifiedAgentic78.7%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld Verified · source release snapshot; version not specifiedAgentic78.7%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld-Verified · source release snapshot; version not specifiedAgentic79%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    PaperBench · source release snapshot; version not specifiedSupporting evidence57.5%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PerceptionBench · source release snapshot; version not specifiedSupporting evidence55.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence28.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official PostTrainBench results; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence39.3%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PostTrainBench Lite · not specifiedSupporting evidence38.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / PostTrainBench Lite / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ProgramBench · source release snapshot; version not specifiedSupporting evidence70.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ResearchRubrics · source release snapshot; version not specifiedSupporting evidence64%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    RSI Index · not specifiedSupporting evidence41.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / RSI Index / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SaaS-Bench · source release snapshot; version not specifiedSupporting evidence43.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SciCode · source release snapshot; version not specifiedCoding56.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    SEC-Bench Pro · May2026 / public graderSupporting evidence45.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; public grader; May2026 JavaScript subset

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SpreadsheetBench 2 · source release snapshot; version not specifiedSupporting evidence29.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SpreadsheetBench v1 · source release snapshot; version not specifiedSupporting evidence88.1%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SVG-Bench · source release snapshot; version not specifiedSupporting evidence58.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-Bench Pro · not specifiedCoding59.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding58.6%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding58.6%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Verified · source release snapshot; version not specifiedCoding82.9%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-fficiency · source release snapshot; version not specifiedSupporting evidence46.6%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-Marathon · source release snapshot; version not specifiedSupporting evidence14%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWEAtlas-QnA · source release snapshot; version not specifiedSupporting evidence45.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWEAtlas-TestWriting · source release snapshot; version not specifiedSupporting evidence42.6%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Terminal-Bench 2.1Coding77.98%Terminal-Bench 2.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-04-23
    Terminal-Bench 2.1 · 2.1Coding85.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding83.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Terminal-Bench 2.1 · 2.1Coding83.4%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding78.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    MiniMax cites this comparator from the official Terminal-Bench leaderboard; a common protocol with its internal Terminus 2 run is not established. See MiniMax M3 benchmark figure, Terminal-bench 2.1 methodology.

    Reviewed 2026-09-06
    Toolathlon · not specifiedAgentic55.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / GPT‑5.5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified · source release snapshot; version not specifiedAgentic73.5%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic73.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    contributes to capability
    lab self-reportReviewed 2026-09-12
    USAMO 2026 · source release snapshot; version not specifiedSupporting evidence98.2%
    Reported settings & source

    MiniMax points out of42; comparator percentages as printed

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    VIBE-V2 · source release snapshot; version not specifiedSupporting evidence50.5%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Video-MME (w. sub) · source release snapshot; version not specifiedSupporting evidence89.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    VideoMME with subtitles · source release snapshot; version not specifiedSupporting evidence89.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    VideoMMMU · source release snapshot; version not specifiedSupporting evidence86.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    WebArena Verified · source release snapshot; version not specifiedAgentic67%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    WorldVQA ForceAnswer · source release snapshot; version not specifiedSupporting evidence38.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    YC-Bench · source release snapshot; version not specifiedSupporting evidence1300000 USD
    Reported settings & source

    Launch figure; agent final assets

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence22%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence41%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    τ³-Banking · source release snapshot; version not specifiedSupporting evidence31.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    GDPval-AAAgentic
    GPQA DiamondHard reasoning
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    OSWorld-VerifiedAgentic
    SWE-bench ProAgentic
    SWE-bench VerifiedAgentic