RankingClaude Fable 5

Claude Fable 5

Data updated 12 Sept 2026

115 published benchmark measures · 24 benchmark families contribute across 6 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Fable 5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Claude Fable 5 · Agentic: 80.2 · SupportedClaude Fable 5 · Hard reasoning: 75.0 · SupportedClaude Fable 5 · Coding: 72.8 · SupportedClaude Fable 5 · Human pref: 64.2 · SupportedClaude Fable 5 · Multimodal: 80.9 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic9 families · 1 with independent evidence · Supported80.2

9 core families; 18 direct opponents across 7 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 76.0–83.4; 0/11 scenarios unsupported. Without one publisher: 74.5–83.1; 0/15 unsupported. Smoothing check: 72.4–84.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning5 families · 2 with independent evidence · Supported75.0

5 core families; 57 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 62.8–79.2; 0/7 scenarios unsupported. Without one publisher: 65.6–76.1; 0/16 unsupported. Smoothing check: 68.5–78.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding4 families · 1 with independent evidence · Supported72.8

4 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 70.3–76.3; 0/7 scenarios unsupported. Without one publisher: 68.3–77.8; 0/20 unsupported. Smoothing check: 67.3–75.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported64.2

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 56.4–74.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 100.0 observed win share
    lmarena.ai · Source 1
Knowledge2 families · 0 with independent evidence · PreliminaryUnknown

2 core families; 2 direct opponents across 1 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Gmmlu: 50.0 observed win share
    Anthropic · Source 1
  • Milu: 50.0 observed win share
    Anthropic · Source 1
Multimodal3 families · 0 with independent evidence · Supported80.9

3 core families; 6 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 77.9–84.8; 2/8 scenarios unsupported. Without one publisher: 80.0–86.5; 1/8 unsupported. Smoothing check: 74.5–84.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Long contextNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 6 effort levels across 73 benchmark/harness combinations →

    Reported effort · Mixed settings

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    24 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Anthropic
    Catalog status
    active
    Availability
    Public provider catalog; account and region restrictions may apply
    Family
    Claude Fable
    Released
    Context
    License
    proprietary
    Model card
    https://platform.claude.com/docs/en/about-claude/model-deprecations
    Default Capability family coverage
    Documented-evidence family coverage/badge/claude-fable-5.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    $OneMillion-Bench (expert score) · source release snapshot; version not specifiedSupporting evidence55.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA Intelligence Index · source release snapshot; version not specifiedSupporting evidence62 index points
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AA-BriefcaseSupporting evidence1572
    Reported settings & source

    Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1583
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-Briefcase (Elo) · source release snapshot; version not specifiedSupporting evidence1574
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AA-LCR · source release snapshot; version not specifiedLong context70%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Agents' Last Exam · not specifiedSupporting evidence48.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · not specifiedSupporting evidence40.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence25.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12

    Published configuration

    Effort: XHigh with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence23.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AndroidBench · source release snapshot; version not specifiedSupporting evidence84.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic43.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    APEX-Agents · source release snapshot; version not specifiedAgentic59.2%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    APEX-SWE · source release snapshot; version not specifiedSupporting evidence58.8%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ARC-AGI · 1Hard reasoning98.5%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI · 2Hard reasoning89.2%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI-1 · 1Hard reasoning98.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-2Hard reasoning89.2%ARC Prize verified

    Effort: Not specified

    contributes to capability
    official board2026-06-09
    ARC-AGI-2 · 2Hard reasoning89.2%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence77.2 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.4 · v1.4Supporting evidence67.2 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence59.9 index score
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence62.1 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Automation-Bench (Pass@1) · source release snapshot; version not specifiedAgentic29.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    AutomationBenchAgentic17.05%
    Reported settings & source

    Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench · not specifiedAgentic17.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · not specifiedAgentic17.4%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · source release snapshot; version not specifiedAgentic29.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench (Public) · source release snapshot; version not specifiedAgentic29.1%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    AutomationBench (v1.0.6) · v1.0.6Agentic46.2%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    BabyVision w/ python · source release snapshot; version not specifiedSupporting evidence90.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BenchCAD · not specifiedSupporting evidence67.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled; three Anthropic evaluation modifications

    Provider-published result; comparator measurements are not automatically independently reproduced. Astra launch footnote5: Claude scores use three modifications described in the Fable5.1 system card; not same controlled setting.

    GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.376 voxel IoU
    Reported settings & source

    Random1000 of17900files; five runs; adaptive thinking max; no tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.675 voxel IoU
    Reported settings & source

    Random1000 of17900files; five runs; adaptive thinking max; with tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BrowseComp · not specifiedAgentic87.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic88%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    ChartographyMultimodal36.6%
    Reported settings & source

    100tasks; adaptive thinking max; five runs; no tools; tools condition has container, standard libraries and crop tool.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ChartographyMultimodal84.2%
    Reported settings & source

    100tasks; adaptive thinking max; five runs; with tools; tools condition has container, standard libraries and crop tool.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal88.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv (RQ) · source release snapshot; version not specifiedMultimodal93.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CorpFin v2 · source release snapshot; version not specifiedSupporting evidence71.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CoWorkBench · source release snapshot; version not specifiedSupporting evidence75.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CritPt · source release snapshot; version not specifiedSupporting evidence28.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CursorBench · 3.2.0Supporting evidence70.5%
    Reported settings & source

    Cursor production agent harness; independently measured by Cursor; max effort.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CursorBench v3.2 · v3.2Supporting evidence70.5%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Cybergym · source release snapshot; version not specifiedSupporting evidence83.1%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    CyberGym · source release snapshot; version not specifiedSupporting evidence83.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DeepSearchQA (F1) · source release snapshot; version not specifiedAgentic94.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSearchQA (F1) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE · source release snapshot; version not specifiedCoding70%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    DeepSWE · source release snapshot; version not specifiedCoding70%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    DeepSWE (v1.1) · v1.1Coding69.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    DeepSWE 1.1 · 1.1Coding70%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    DeepSWE v1.1Coding69.9%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding69.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding69.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1 · v1.1Coding70%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DSBench-FullStack † · source release snapshot; version not specifiedSupporting evidence77.2%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result. † source footnote applies.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DSBench-Hard † · source release snapshot; version not specifiedSupporting evidence68.3%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result. † source footnote applies.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitBench · source release snapshot; version not specifiedSupporting evidence78%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence181 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence247 count
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence56.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierCode · 1.1 ExtendedSupporting evidence64.9%
    Reported settings & source

    Cognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierCode · 1.1 MainSupporting evidence53.5%
    Reported settings & source

    Cognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12

    Published configuration

    Effort: XHigh

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierCode 1.1 Extended (score) · 1.1Supporting evidence64.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode 1.1 Main (score) · 1.1Supporting evidence53.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode v1.1 Extended · v1.1Supporting evidence63.6%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierMath Tier 1-3 (v2) · v2Hard reasoning87%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning90.2%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning87.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierSWE · 2Supporting evidence0.48 fraction
    Reported settings & source

    Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence88.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence86.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    FrontierSWE · source release snapshot; version not specifiedSupporting evidence88.2%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    Externally evaluated result, attributed in the source footnote.

    Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    gdp.pdf · not specifiedSupporting evidence29.8%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval-AA · 2Agentic1723
    Reported settings & source

    Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GDPval-AA v2 · source release snapshot; version not specifiedAgentic1743
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    Externally evaluated result, attributed in the source footnote.

    Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    GDPval-AA v2 · v2Agentic1760
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1747
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    GDPval-AA v2 (Elo) · source release snapshot; version not specifiedAgentic1741
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GMMLUKnowledge93.6%
    Reported settings & source

    Mean accuracy42languages; adaptive max; one trial; no tools/custom system prompts.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GPQA DiamondHard reasoning92.626%GPQA Diamond reported

    Effort: Not specified

    contributes to capability
    independent repro2026-09-11
    GPQA Diamond · not specifiedHard reasoning92.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · not specifiedHard reasoning92.6%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning92.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning92.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Harvey LAB (Vals) · source release snapshot; version not specifiedSupporting evidence11.3%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Harvey Lab-AA · source release snapshot; version not specifiedSupporting evidence93.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBenchSupporting evidence61.2%
    Reported settings & source

    Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench ProfessionalSupporting evidence68.9%
    Reported settings & source

    Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench ProfessionalSupporting evidence63.3%
    Reported settings & source

    Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench Professional · not specifiedSupporting evidence60.9%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBench Professional (length-adjusted) · not specifiedSupporting evidence60.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted; OpenAI reproduction; GPT-5.4 grader; unclipped

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE · source release snapshot; version not specifiedHard reasoning53.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning53.3%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning63%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    HLE w/ tools · source release snapshot; version not specifiedHard reasoning64.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    HLE w/ Tools · source release snapshot; version not specifiedHard reasoning63.9%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning53.3%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    HLE-Full · source release snapshot; version not specifiedHard reasoning63%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Humanity's Last ExamHard reasoning57.8%HLE no tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-09-01
    Humanity's Last ExamHard reasoning63.8%HLE with tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-09-01
    Humanity’s Last ExamHard reasoning57.8%
    Reported settings & source

    Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

    Published configuration

    Effort: Auto thinking

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Humanity’s Last ExamHard reasoning63.8%
    Reported settings & source

    Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

    Published configuration

    Effort: Auto thinking

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Humanity's Last Exam (w/ tools) · not specifiedHard reasoning63.8%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IFBench · source release snapshot; version not specifiedSupporting evidence63.5%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Internal Data Science Tasks · not specifiedSupporting evidence34.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Internal Data Science Tasks / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Database Migration Tasks · not specifiedSupporting evidence50.3%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Internal Design Tasks · not specifiedSupporting evidence35.8%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Internal Design Tasks / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    JobBench · source release snapshot; version not specifiedSupporting evidence57.4%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    JobBench · source release snapshot; version not specifiedSupporting evidence57.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Kimi Code Bench 2.0 · 2.0Supporting evidence76.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Legal Research Bench · source release snapshot; version not specifiedSupporting evidence49.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    LMArena Text ArenaHuman pref1506LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    Management Consulting Tasks (Internal) · not specifiedSupporting evidence35.5%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    MathVision · source release snapshot; version not specifiedSupporting evidence94.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MathVision · source release snapshot; version not specifiedSupporting evidence98.6%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MCP-Atlas · source release snapshot; version not specifiedAgentic84.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MCPMark-Verified · source release snapshot; version not specifiedSupporting evidence87.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MILUKnowledge92.9%
    Reported settings & source

    Mean accuracy11languages; adaptive max; five trials; no tools/custom system prompts.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence49.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence49.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal81.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    MMMU-Pro · source release snapshot; version not specifiedMultimodal86.5%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OfficeQA ProAgentic57.9%
    Reported settings & source

    Databricks evaluation reading documents as images; differs from extracted-text harness.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-12
    OfficeQA Pro · source release snapshot; version not specifiedAgentic69.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OmniDocBench · source release snapshot; version not specifiedSupporting evidence89.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OSWorld · 2.0 August2026 task releaseAgentic72.9%
    Reported settings & source

    partial pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

    Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld · 2.0 August2026 task releaseAgentic36.1%
    Reported settings & source

    strict pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

    Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld 2.0 · 2.0Agentic66.1%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld-VerifiedAgentic85.96%OSWorld-Verified reported

    Effort: Not specified

    contributes to capability
    official board2026-08-01
    OSWorld-Verified · source release snapshot; version not specifiedAgentic85%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    PaperBench · source release snapshot; version not specifiedSupporting evidence88.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PerceptionBench · source release snapshot; version not specifiedSupporting evidence57.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PLawBench · source release snapshot; version not specifiedSupporting evidence70.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence41.4%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PostTrainBench · source release snapshot; version not specifiedSupporting evidence41.8%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PRBench-Finance · source release snapshot; version not specifiedSupporting evidence55.8%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    PRBench-Legal · source release snapshot; version not specifiedSupporting evidence57.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench · 166 golden-task subsetSupporting evidence86.3%
    Reported settings & source

    mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

    Hidden-test pass rate, not fraction of completely solved programs. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench · source release snapshot; version not specifiedSupporting evidence76.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProgramBench (Almost Solved) · source release snapshot; version not specifiedSupporting evidence33%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenQoderBench · source release snapshot; version not specifiedSupporting evidence63.1%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenReactBench · source release snapshot; version not specifiedSupporting evidence1770
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSVGBench · source release snapshot; version not specifiedSupporting evidence1690
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    QwenSWEBench · source release snapshot; version not specifiedSupporting evidence86.3%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SciCode · source release snapshot; version not specifiedCoding60.2%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    SkillsBench · source release snapshot; version not specifiedSupporting evidence70.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SpreadsheetBench 2 · source release snapshot; version not specifiedSupporting evidence34.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-bench MultilingualCoding86.6%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-bench MultimodalSupporting evidence54.1%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-bench ProCoding80%Anthropic SWE-bench Pro reported system

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-09-01
    SWE-bench ProCoding80%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-Bench Pro · not specifiedCoding80%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding80%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    SWE-Marathon · source release snapshot; version not specifiedSupporting evidence35%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-Marathon (v1.1) · v1.1Supporting evidence33.1%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding84.6%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses.

    Cited comparator result; see the benchmark footnote.

    Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding88%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    Terminal Bench 2.1 · 2.1Coding88%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    Terminal Bench 3.0 · 3.0Coding33.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    Terminal-Bench · 4.0Coding42%
    Reported settings & source

    Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

    Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal-Bench 2.1 · 2.1Coding83.1%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding88%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison.

    Reviewed 2026-09-12
    Terminal-Bench 4.0 · 4.0Coding44.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench Science 0.1 · not specifiedHard reasoning21.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench v3.0 · v3.0Coding34.1%
    Reported settings & source

    Grok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI.

    First-party reported result; comparator results retain the source evaluation setup.

    Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench-Science · 0.1Coding24.7%
    Reported settings & source

    70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board.

    Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon · not specifiedAgentic61.7%
    Reported settings & source

    Launch table reported configuration; per-cell reasoning effort unspecified

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / Claude Fable 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified · source release snapshot; version not specifiedAgentic74.7%
    Reported settings & source

    GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

    First-party reported result.

    Comparison limit: The source column includes fallback execution without a documented exact effort.

    zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source column includes fallback execution without a documented exact effort.

    Reviewed 2026-09-12
    Toolathlon Verified (Pass@1) · source release snapshot; version not specifiedAgentic77.9%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Reviewed 2026-09-12
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic77.9%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon-Verified · source release snapshot; version not specifiedAgentic77.9%
    Reported settings & source

    DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

    First-party reported result.

    Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs.

    Reviewed 2026-09-12
    WideSearch · source release snapshot; version not specifiedSupporting evidence81.2%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Item-F1 over4runs; Qwen-Agent for Qwen,ClaudeCode for comparators.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: WideSearch · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WorkSpaceBench · source release snapshot; version not specifiedSupporting evidence68.7%
    Reported settings & source

    Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

    First-party reported result.

    Comparison limit: The source states that Fable 5 results may involve fallback execution; an exact configuration is not established.

    Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    WorldVQA ForceAnswer · source release snapshot; version not specifiedSupporting evidence56.7%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence23%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ZeroBench (pass@5) · source release snapshot; version not specifiedSupporting evidence46%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    First-party reported result.

    moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    τ³-Banking · source release snapshot; version not specifiedSupporting evidence26.8%
    Reported settings & source

    Table headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs.

    Cited result; see benchmark-specific Evaluation Details.

    Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison.

    moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12

    Published configuration

    Effort: Max with fallbacks

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    GDPval-AAAgentic
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    SWE-bench VerifiedAgentic
    Terminal-Bench 2.1Agentic