RankingGPT-5.6 Sol

GPT-5.6 Sol

Data updated 12 Sept 2026

153 published benchmark measures · 25 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-5.6 Sol capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-5.6 Sol · Agentic: 66.0 · SupportedGPT-5.6 Sol · Hard reasoning: 75.9 · SupportedGPT-5.6 Sol · Coding: 72.0 · SupportedGPT-5.6 Sol · Human pref: 51.0 · SupportedGPT-5.6 Sol · Multimodal: 69.7 · SupportedGPT-5.6 Sol · Long context: 70.0 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic9 families · 1 with independent evidence · Supported66.0

9 core families; 56 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 61.8–69.3; 0/11 scenarios unsupported. Without one publisher: 62.4–68.2; 0/15 unsupported. Smoothing check: 61.4–68.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning5 families · 2 with independent evidence · Supported75.9

5 core families; 60 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 70.4–78.9; 0/7 scenarios unsupported. Without one publisher: 71.1–76.6; 0/16 unsupported. Smoothing check: 69.4–79.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding3 families · 1 with independent evidence · Supported72.0

3 core families; 30 direct opponents across 9 labs.

Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 62.0–80.6; 0/7 scenarios unsupported. Without one publisher: 70.4–78.0; 0/20 unsupported. Smoothing check: 67.0–74.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported51.0

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 50.5–51.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 93.9 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal4 families · 0 with independent evidence · Supported69.7

    4 core families; 12 direct opponents across 4 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 69.0–79.0; 0/8 scenarios unsupported. Without one publisher: 49.8–75.4; 0/8 unsupported. Smoothing check: 65.0–72.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context3 families · 0 with independent evidence · Supported70.0

    3 core families; 8 direct opponents across 4 labs.

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: 65.9–79.1; 1/5 scenarios unsupported. Without one publisher: 70.0–75.5; 2/5 unsupported. Smoothing check: 68.1–70.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 7 effort levels across 72 benchmark/harness combinations →

    Reported effort · Mixed settings

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    25 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    OpenAI
    Catalog status
    active
    Availability
    Public provider catalog; account and region restrictions may apply
    Family
    GPT-5.6
    Released
    Context
    1,050,000 tokens
    License
    proprietary
    Model card
    https://developers.openai.com/api/docs/models/all
    Default Capability family coverage
    Documented-evidence family coverage/badge/gpt-5.6-sol.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    $OneMillion-Bench (expert score) · source release snapshot; version not specifiedSupporting evidence53.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA Intelligence Index · source release snapshot; version not specifiedSupporting evidence61 index points
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AA-BriefcaseSupporting evidence1502
    Reported settings & source

    Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1495
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1502
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AA-LCR · source release snapshot; version not specifiedLong context73.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Agentic IF Index · internalSupporting evidence60.5 index
    Reported settings & source

    Internal composite instruction-following evaluations; no fixed task count; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · not specifiedSupporting evidence53.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · not specifiedSupporting evidence52.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence29.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence28.6%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence30.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence53.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AndroidBench · source release snapshot; version not specifiedSupporting evidence74%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic39.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic56.7%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI · 1Hard reasoning96.5%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI · 2Hard reasoning92.5%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI-1 · 1Hard reasoning97.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-2Hard reasoning92.5%ARC Prize verified

    Effort: Not specified

    contributes to capability
    official board2026-07-09
    ARC-AGI-2 · 2Hard reasoning92.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-3 · 3Hard reasoning7.8%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-3 · 3Hard reasoning7.78%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence80 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.4 · v1.4Supporting evidence65.1 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence IndexSupporting evidence61 AA Intelligence Index

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-08-12
    Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence58.9 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence60.9 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Automation-Bench (Pass@1) · source release snapshot; version not specifiedAgentic29.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBenchAgentic19.6%
    Reported settings & source

    Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench · not specifiedAgentic18.1%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · not specifiedAgentic18.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · public v3Agentic46.7%
    Reported settings & source

    600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · source release snapshot; version not specifiedAgentic29.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench (v1.0.6) · v1.0.6Agentic45.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    BabyVision w/ python · source release snapshot; version not specifiedSupporting evidence88.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BenchCAD · not specifiedSupporting evidence83.3%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD · not specifiedSupporting evidence70.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD (python tool) · not specifiedSupporting evidence83.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Big Finance Bench · not specifiedSupporting evidence53%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BioMysteryBench · Human DifficultSupporting evidence28.8%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BioMysteryBench · Human DifficultSupporting evidence44.7%
    Reported settings & source

    Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BioMysteryBench · Human SolvableSupporting evidence86.1%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BioMysteryBench · Human SolvableSupporting evidence79.5%
    Reported settings & source

    Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseCompAgentic90.4%BrowseComp reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-09
    BrowseComp · not specifiedAgentic90.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · not specifiedAgentic90.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · not specifiedAgentic92.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Sol Ultra · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic90.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Capture-the-Flag Challenges · not specifiedSupporting evidence96.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / Capture-the-Flag Challenges / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal84.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal89.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv ReasoningMultimodal85.8%
    Reported settings & source

    No tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Coronavirus-ACE2 Cell-Entry Screen · not specifiedSupporting evidence0.43 composite score
    Reported settings & source

    System-card reported configuration; reasoning effort unspecified

    Observed named-model performance; helpful-only checkpoint omitted as distinct noncatalog identity.

    GPT-6 Astra System Card · Section10.1.1.2.3 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CorpFin v2 · source release snapshot; version not specifiedSupporting evidence64.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CoWorkBench · source release snapshot; version not specifiedSupporting evidence71.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CritPt · source release snapshot; version not specifiedSupporting evidence32.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CursorBench · 3.2.0Supporting evidence67.2%
    Reported settings & source

    Cursor production agent harness; independently measured by Cursor; max effort.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CursorBench v3.2 · v3.2Supporting evidence67.2%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CyberGym · source release snapshot; version not specifiedSupporting evidence83.6%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DeepSearchQAAgentic93.1 percent F1
    Reported settings & source

    900questions; common search backend/browser harness; answer-set F1; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE · 1.1Coding72.7%
    Reported settings & source

    Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

    Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE · 1.1Coding73%
    Reported settings & source

    113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max

    Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE · source release snapshot; version not specifiedCoding73%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE (v1.1) · v1.1Coding72.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE 1.1 · 1.1Coding73%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE v1.1Coding72.7%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding72.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding72.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding73%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ExploitBench · not specifiedSupporting evidence78.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitBench · not specifiedSupporting evidence73.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitBench · source release snapshot; version not specifiedSupporting evidence76.5%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitBench (June-Aug 2026) · June-Aug2026Supporting evidence5.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; 20 vulnerabilities / 13 Chrome releases; 300-turn limit

    Provider-published result; comparator measurements are not automatically independently reproduced. Footnote14 separately reports11.5% with fewer turn-limit interruptions; main table remains5.5%.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench (June-Aug 2026) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitBench (June-Aug 2026) · June-Aug2026Supporting evidence11.5%
    Reported settings & source

    Similar settings with fewer300-turn-limit interruptions; production safeguards absent

    Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

    GPT-6 Astra: A new generation of intelligence · Footnote14 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym · not specifiedSupporting evidence30.3%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; v1 offline environment; no runtime package installation; token capped, no wall-clock cap

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym · not specifiedSupporting evidence33.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; six-hour evaluation cap; alpha API latency rescaled to public API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitGym / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym · not specifiedSupporting evidence24.9%
    Reported settings & source

    Two-hour cap; alpha API latency rescaled; reduced safeguards

    Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

    GPT-5.6: Frontier intelligence that scales with your ambition · Pushing the frontier on cyber and science paragraph · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence216 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence293 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence53.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence53.8%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode 1.1 Extended (score) · 1.1Supporting evidence60.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode 1.1 Main (score) · 1.1Supporting evidence47.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode v1.1 Extended · v1.1Supporting evidence60.6%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierMath Tier 1-3 (v2) · v2Hard reasoning89%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning83%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning83%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierSWE · 2Supporting evidence0.32 fraction
    Reported settings & source

    Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence71.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    GDP.PDFSupporting evidence40%
    Reported settings & source

    All-pass rate; allmodels selfcomputed byGoogle

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    gdp.pdf · not specifiedSupporting evidence30.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval-AAAgentic1624Artificial Analysis GDPval-AA

    Effort: Not specified

    contributes to capability
    official board2026-09-12
    GDPval-AA · 2Agentic1711
    Reported settings & source

    Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GDPval-AA · 2Agentic1710
    Reported settings & source

    Artificial Analysis publicboard snapshot; effort as reported

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA · 2Agentic1710
    Reported settings & source

    Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 · source release snapshot; version not specifiedAgentic1730
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    Externally evaluated result, attributed in the source footnote.

    Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

    zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 · v2Agentic1748
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1736
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1728
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GeneBench Pro · not specifiedSupporting evidence28.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GeneBench Pro · v13Supporting evidence32.3%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / GeneBench Pro / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GPQA DiamondHard reasoning94.141%GPQA Diamond reported

    Effort: Not specified

    contributes to capability
    independent repro2026-09-11
    GPQA Diamond · not specifiedHard reasoning94.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · not specifiedHard reasoning94.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning94.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning94.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GraphWalks BFS 1mil f1 · not specifiedLong context77.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GraphWalks BFS 256k f1 · not specifiedLong context90.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Harvey LAB (Vals) · source release snapshot; version not specifiedSupporting evidence2.5%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Harvey Lab-AA · source release snapshot; version not specifiedSupporting evidence87.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Harvey Legal Agent Benchmark · source release snapshot; version not specifiedSupporting evidence2.5%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench · not specifiedSupporting evidence57 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 1764 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench · not specifiedSupporting evidence55.6 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 1764 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-sol / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench · source release snapshot; version not specifiedSupporting evidence55.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench Consensus · not specifiedSupporting evidence95.5 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 1740 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Consensus · not specifiedSupporting evidence95.3 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 1740 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-sol / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Hard · not specifiedSupporting evidence33.1 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 1751 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Hard · not specifiedSupporting evidence31.1 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 1751 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-sol / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence60.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence60.5 score (0-100)
    Reported settings & source

    length-adjusted; official HealthBench scoring

    Mean response length 3228 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional · not specifiedSupporting evidence64.1 score (0-100)
    Reported settings & source

    unadjusted; official HealthBench scoring

    Mean response length 3228 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

    GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-sol / unadjusted · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional (length-adjusted) · not specifiedSupporting evidence60.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE · source release snapshot; version not specifiedHard reasoning47.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE w/ tools · source release snapshot; version not specifiedHard reasoning58%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE w/ Tools · source release snapshot; version not specifiedHard reasoning64.5%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning44.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning58%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Verified · source release snapshot; version not specifiedHard reasoning54.5%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IFBench · source release snapshot; version not specifiedSupporting evidence72.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Internal Data Science Tasks · not specifiedSupporting evidence30.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Internal Data Science Tasks / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Database Migration Tasks · not specifiedSupporting evidence42.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Design Tasks · not specifiedSupporting evidence47.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Internal Design Tasks / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Research Debugging Evaluation · not specifiedSupporting evidence68.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / Internal Research Debugging Evaluation / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    JobBenchSupporting evidence45.4%
    Reported settings & source

    65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    JobBench · source release snapshot; version not specifiedSupporting evidence45.4%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence45.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    KernelGen 1P · not specifiedSupporting evidence61.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / KernelGen 1P / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Kimi Code Bench 2.0 · 2.0Supporting evidence64.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    LABBench · 2Supporting evidence82.1%
    Reported settings & source

    Selfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Legal Research Bench · source release snapshot; version not specifiedSupporting evidence48.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    LifeSciBench · Gold v1Supporting evidence59.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / LifeSciBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LifeSciBench · not specifiedSupporting evidence59.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LMArena Text ArenaHuman pref1482LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    LongBench v2 · source release snapshot; version not specifiedLong context67.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    LVBench · staticMultimodal82.1%
    Reported settings & source

    No tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024

    Frame budgets differ; table labels Gemini3.8static explicitly.

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Management Consulting Tasks (Internal) · not specifiedSupporting evidence43.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MathVision · source release snapshot; version not specifiedSupporting evidence95.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MathVision · source release snapshot; version not specifiedSupporting evidence97.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MCP-Atlas · source release snapshot; version not specifiedAgentic83.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MCPMark-Verified · source release snapshot; version not specifiedSupporting evidence92.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MedChemBench (Internal) · not specifiedSupporting evidence47.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MedChemBench (Internal) · not specifiedSupporting evidence48.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence46.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence46.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MMMU Pro (no tools) · not specifiedMultimodal83%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; no tools

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMMU Pro (with tools) · not specifiedMultimodal84.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; tools enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (with tools) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMMU-Pro · source release snapshot; version not specifiedMultimodal83%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal84.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMVU · source release snapshot; version not specifiedSupporting evidence81.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MRCR · v2 256K-512KLong context91.5 percent sequence match
    Reported settings & source

    8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MRCR · v2 512K-1MLong context73.8 percent sequence match
    Reported settings & source

    8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MRCR v2 256K (8-needle) · source release snapshot; version not specifiedLong context93.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    NanoGPT · not specifiedSupporting evidence9.69%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / NanoGPT / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    No-CoT math time horizon · not specifiedSupporting evidence3.6 minutes
    Reported settings & source

    UK AISI; single forward pass; no chain of thought

    Time-horizon estimate; possible contamination noted by evaluator. Higher means harder tasks solved, not slower inference.

    GPT-6 Astra System Card · Section9.3 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OfficeQA Pro · source release snapshot; version not specifiedAgentic63.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OmniDocBench · source release snapshot; version not specifiedSupporting evidence85.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OpenAI MRCR v2 8-needle 256K-512K · v2Long context91.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 256K-512K

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OpenAI MRCR v2 8-needle 256K-512K · v2Long context91.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 256K-512K

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OpenAI MRCR v2 8-needle 512K-1M · v2Long context73.8%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 512K-1M

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OpenAI MRCR v2 8-needle 512K-1M · v2Long context73.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 512K-1M

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OpenScore String Quartets (1 - OMR-NED) · not specifiedSupporting evidence0.19 1 - OMR-NED
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / OpenScore String Quartets (1 - OMR-NED) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Organic Chemistry · 2 revisedSupporting evidence43.2%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OSWorld 2.0 · 2.0Agentic62.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 · 2.0Agentic62.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic65.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld binary · 2.0 08.08Agentic27.3%
    Reported settings & source

    108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

    Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld partial · 2.0 08.08Agentic62.7%
    Reported settings & source

    108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

    Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld partial · 2.0 task revision unverified gpt-5.6-solAgentic62.6%
    Reported settings & source

    Partialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness

    Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports.

    Comparison limit: Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison.

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison.

    Reviewed 2026-09-06
    OSWorld-Verified · source release snapshot; version not specifiedAgentic83%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    PaperBench · source release snapshot; version not specifiedSupporting evidence90.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PerceptionBench · source release snapshot; version not specifiedSupporting evidence59.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PLawBench · source release snapshot; version not specifiedSupporting evidence72.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence34.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence36.2%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench Lite · not specifiedSupporting evidence50.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / PostTrainBench Lite / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PRBench-Finance · source release snapshot; version not specifiedSupporting evidence55.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PRBench-Legal · source release snapshot; version not specifiedSupporting evidence57.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench · source release snapshot; version not specifiedSupporting evidence77.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench (Almost Solved) · source release snapshot; version not specifiedSupporting evidence23%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProteinGym · HardSupporting evidence35.5 percent rank correlation
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protocols · TroubleshootingSupporting evidence56.4%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protocols · Understanding network-restrictedSupporting evidence63.9%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenQoderBench · source release snapshot; version not specifiedSupporting evidence53.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenReactBench · source release snapshot; version not specifiedSupporting evidence1564
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSVGBench · source release snapshot; version not specifiedSupporting evidence1758
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSWEBench · source release snapshot; version not specifiedSupporting evidence73.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ResearchRubrics · source release snapshot; version not specifiedSupporting evidence73.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    RSI Index · not specifiedSupporting evidence57.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / RSI Index / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SaaS-Bench · source release snapshot; version not specifiedSupporting evidence61.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Sandbox Bench · September2026 internalSupporting evidence4.5%
    Reported settings & source

    22 isolated CTF-style targets; protected-flag success metric

    Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

    GPT-6 Astra System Card · Section10.1.2.4 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SciCode · source release snapshot; version not specifiedCoding56.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    ScreenSpot-Pro (no tools) · not specifiedMultimodal76.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; no tools

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / ScreenSpot-Pro (no tools) / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SEC-Bench Pro · May2026 / public graderSupporting evidence71.2%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; public grader; May2026 JavaScript subset

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SEC-Bench Pro · May2026 / public graderSupporting evidence74.3%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration; reduced or absent production safeguards; public grader; May2026 JavaScript subset

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol Ultra · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SEC-Bench Pro · May2026 / revised root-cause graderSupporting evidence79.1%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; May2026 JavaScript subset,183 vulnerabilities; revised agent root-cause grader

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SHP2 Protein Function Prediction · 3 unpublished assay datasetsSupporting evidence0.3 mean R-squared
    Reported settings & source

    System-card reported configuration; reasoning effort unspecified

    Production-named model result; separate helpful-only checkpoint omitted because no exact catalog identity.

    GPT-6 Astra System Card · Section10.1.1.2.2 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SkillsBench · source release snapshot; version not specifiedSupporting evidence73.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SpreadsheetBench 2 · source release snapshot; version not specifiedSupporting evidence32.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SRE-Bench · 262 binaries / 19 programsSupporting evidence55.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SRE-Bench · 262 binaries / 19 programsSupporting evidence68.7%
    Reported settings & source

    pass@4; four independent trials; all six objectives required; reduced production safeguards

    Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

    GPT-6 Astra System Card · Section10.1.2.3 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-Atlas Codebase QnASupporting evidence53.5%
    Reported settings & source

    124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-bench ProCoding64.6%SWE-bench Pro reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-09
    SWE-bench ProCoding64.6%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-Bench Pro · not specifiedCoding64.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding64.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-Marathon · source release snapshot; version not specifiedSupporting evidence39%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-Marathon (v1.1) · v1.1Supporting evidence42.5%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding88.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding88.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal Bench 3.0 · 3.0Coding34.6%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal-Bench · 2.1Coding88.8%
    Reported settings & source

    Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench · 2.1Coding88.8%
    Reported settings & source

    89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: OpenAI (exact harness revision not specified)

    Native harnesses differ; not a Terminus2-only comparison.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Different native coding harnesses are mixed; exact shared agent configuration is not established.

    Reviewed 2026-09-06
    Terminal-Bench · 4.0Coding37.3%
    Reported settings & source

    Claude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board.

    Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed.

    Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison.

    Reviewed 2026-09-12
    Terminal-Bench · 4.0Coding37.3%
    Reported settings & source

    Officialpublicboard highest scoring thinking level; nativeagents may differ

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1Coding88.8%Terminal-Bench 2.1 reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-09
    Terminal-Bench 2.1Coding91.9%Codex ultra

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-09
    Terminal-Bench 2.1 · 2.1Coding88.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding91.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Sol Ultra · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding88.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Terminal-Bench 4.0 · 4.0Coding37.3%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench Science 0.1 · not specifiedHard reasoning22.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Best across efforts

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench v3.0 · v3.0Coding34.6%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench-Science · 0.1Coding22.4%
    Reported settings & source

    70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon · not specifiedAgentic58%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / GPT‑5.6 Sol · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified · source release snapshot; version not specifiedAgentic74.9%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon Verified (Pass@1) · source release snapshot; version not specifiedAgentic74.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic74.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Video-MME (w. sub) · source release snapshot; version not specifiedSupporting evidence89.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    VulcanBench v3Supporting evidence87%VulcanBench v3 bare-bones API · Report 07 · high

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-12
    VulcanBench v3Supporting evidence78.3%VulcanBench v3 bare-bones API · Report 07 · low

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-12
    VulcanBench v3Supporting evidence82.6%VulcanBench v3 bare-bones API · Report 07 · medium

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-12
    WorkSpaceBench · source release snapshot; version not specifiedSupporting evidence65.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WorldVQA ForceAnswer · source release snapshot; version not specifiedSupporting evidence41.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence17%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence35%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    τ³-Banking · source release snapshot; version not specifiedSupporting evidence33%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    OSWorld-VerifiedAgentic
    SWE-bench VerifiedAgentic