RankingClaude Fable 5.1

Claude Fable 5.1

Data updated 12 Sept 2026

40 published benchmark measures · 16 benchmark families contribute across 5 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Fable 5.1 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Claude Fable 5.1 · Agentic: 97.8 · SupportedClaude Fable 5.1 · Hard reasoning: 86.8 · SupportedClaude Fable 5.1 · Coding: 89.9 · SupportedClaude Fable 5.1 · Multimodal: 95.1 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic4 families · 0 with independent evidence · Supported97.8

4 core families; 4 direct opponents across 2 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 97.1–98.9; 2/11 scenarios unsupported. Without one publisher: 96.8–98.7; 1/15 unsupported. Smoothing check: 92.0–99.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning5 families · 2 with independent evidence · Supported86.8

5 core families; 57 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 80.7–90.7; 0/7 scenarios unsupported. Without one publisher: 69.6–88.7; 0/16 unsupported. Smoothing check: 79.8–89.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding4 families · 0 with independent evidence · Supported89.9

4 core families; 5 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 87.2–93.4; 0/7 scenarios unsupported. Without one publisher: 74.1–91.8; 0/20 unsupported. Smoothing check: 82.6–93.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Knowledge2 families · 0 with independent evidence · PreliminaryUnknown

    2 core families; 2 direct opponents across 1 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Gmmlu: 100.0 observed win share
      Anthropic · Source 1
    • Milu: 100.0 observed win share
      Anthropic · Source 1
    Multimodal1 families · 0 with independent evidence · Preliminary95.1

    1 core families; 2 direct opponents across 1 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long contextNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 5 effort levels across 63 benchmark/harness combinations →

      Reported effort · Mixed settings

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      16 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      Anthropic
      Catalog status
      active
      Availability
      Public provider catalog; account and region restrictions may apply
      Family
      Claude Fable
      Released
      Context
      1,000,000 tokens
      License
      proprietary
      Model card
      https://platform.claude.com/docs/en/about-claude/model-deprecations
      Default Capability family coverage
      Documented-evidence family coverage/badge/claude-fable-5-1.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      AA-BriefcaseSupporting evidence1694
      Reported settings & source

      Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-BriefcaseSupporting evidence1686
      Reported settings & source

      Artificial Analysis; xhigh effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-BriefcaseSupporting evidence1611
      Reported settings & source

      Artificial Analysis; high effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: High

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-BriefcaseSupporting evidence61.5%
      Reported settings & source

      Artificial Analysis; max effort; component rubric pass.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-BriefcaseSupporting evidence2025 rating
      Reported settings & source

      Artificial Analysis; max effort; component analytical quality.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      AA-BriefcaseSupporting evidence1495 rating
      Reported settings & source

      Artificial Analysis; max effort; component presentation.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ARC-AGI · 1Hard reasoning97.5%
      Reported settings & source

      ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      ARC-AGI · 2Hard reasoning90%
      Reported settings & source

      ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      ARC-AGI-1 · 1Hard reasoning97.5%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      ARC-AGI-2Hard reasoning90%ARC Prize verified

      Effort: Not specified

      contributes to capability
      official board2026-09-01
      ARC-AGI-2 · 2Hard reasoning90%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence65.7 index score
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      AutomationBenchAgentic31.4%
      Reported settings & source

      Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      AutomationBench · not specifiedAgentic31.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      BenchCAD · not specifiedSupporting evidence84.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled; three Anthropic evaluation modifications

      Provider-published result; comparator measurements are not automatically independently reproduced. Astra launch footnote5: Claude scores use three modifications described in the Fable5.1 system card; not same controlled setting.

      GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.437 voxel IoU
      Reported settings & source

      Random1000 of17900files; five runs; adaptive thinking max; no tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.843 voxel IoU
      Reported settings & source

      Random1000 of17900files; five runs; adaptive thinking max; with tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      ChartographyMultimodal42.6%
      Reported settings & source

      100tasks; adaptive thinking max; five runs; no tools; tools condition has container, standard libraries and crop tool.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      ChartographyMultimodal86.2%
      Reported settings & source

      100tasks; adaptive thinking max; five runs; with tools; tools condition has container, standard libraries and crop tool.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      CursorBench · 3.2.0Supporting evidence73.4%
      Reported settings & source

      Cursor production agent harness; independently measured by Cursor; max effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      CursorBench · 3.2.0Supporting evidence68%
      Reported settings & source

      Cursor production agent harness; medium effort; $3.53/task.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

      Published configuration

      Effort: Medium

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      DeepSWE · 1.1Coding67.4%
      Reported settings & source

      Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. 113 tasks; original hidden-test grading.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.3 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-12
      DeepSWE v1.1Coding67.4%Anthropic DeepSWE1.1 reported system

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-09-01
      DeepSWE v1.1 · v1.1Coding67.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      FrontierCode · 1.1 ExtendedSupporting evidence63.6%
      Reported settings & source

      Cognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12

      Published configuration

      Effort: Medium

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      FrontierCode · 1.1 MainSupporting evidence50.9%
      Reported settings & source

      Cognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12

      Published configuration

      Effort: Medium

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      FrontierCode 1.1 Extended (score) · 1.1Supporting evidence63.6%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierCode 1.1 Main (score) · 1.1Supporting evidence50.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierMath Tier 4 (v2) · v2Hard reasoning87.8%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      FrontierSWE · 2Supporting evidence0.57 fraction
      Reported settings & source

      Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      GDPval-AA · 2Agentic1853
      Reported settings & source

      Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      GDPval-AA · 2Agentic1835
      Reported settings & source

      Artificial Analysis; xhigh effort; same release board snapshot.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.3 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-12
      GMMLUKnowledge94%
      Reported settings & source

      Mean accuracy42languages; adaptive max; one trial; no tools/custom system prompts.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      GPQA DiamondHard reasoning93.737%GPQA Diamond reported

      Effort: Not specified

      contributes to capability
      independent repro2026-09-11
      GPQA Diamond · not specifiedHard reasoning93.7%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      HealthBenchSupporting evidence66.7%
      Reported settings & source

      Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HealthBenchSupporting evidence60%
      Reported settings & source

      Length-adjusted score using GPT5.5card method; otherwise raw evaluation configuration.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HealthBench ProfessionalSupporting evidence74.2%
      Reported settings & source

      Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HealthBench ProfessionalSupporting evidence62.1%
      Reported settings & source

      Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      HealthBench Professional (length-adjusted) · not specifiedSupporting evidence58.1%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted; OpenAI reproduction; GPT-5.4 grader; unclipped; Opus5 fallback for provider refusals

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Humanity's Last ExamHard reasoning60.9%HLE no tools

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-09-01
      Humanity's Last ExamHard reasoning65%HLE with tools

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-09-01
      Humanity’s Last ExamHard reasoning60.9%
      Reported settings & source

      Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

      Published configuration

      Effort: Auto thinking

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Humanity’s Last ExamHard reasoning65%
      Reported settings & source

      Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

      Published configuration

      Effort: Auto thinking

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Humanity's Last Exam (w/ tools) · not specifiedHard reasoning65%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Internal Database Migration Tasks · not specifiedSupporting evidence57.8%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Legal Agent Benchmark · 120-task held-out subsetSupporting evidence16.7%
      Reported settings & source

      all-pass; Artificial Analysis harness; xhigh effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Legal Agent Benchmark · 120-task held-out subsetSupporting evidence93.3%
      Reported settings & source

      criterion-pass; Artificial Analysis harness; xhigh effort.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

      Published configuration

      Effort: XHigh

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Legal Agent Benchmark · 1235-task public subsetSupporting evidence19.09%
      Reported settings & source

      all-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      Legal Agent Benchmark · 1235-task public subsetSupporting evidence90.81%
      Reported settings & source

      criterion-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      MILUKnowledge93%
      Reported settings & source

      Mean accuracy11languages; adaptive max; five trials; no tools/custom system prompts.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.2 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OfficeQAAgentic80.2%
      Reported settings & source

      Extracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OfficeQA ProAgentic69%
      Reported settings & source

      Extracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OSWorld · 2.0 August2026 task releaseAgentic77.9%
      Reported settings & source

      partial pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

      Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      OSWorld · 2.0 August2026 task releaseAgentic41.7%
      Reported settings & source

      strict pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

      Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      ProgramBench · 166 golden-task subsetSupporting evidence87.6%
      Reported settings & source

      mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

      Hidden-test pass rate, not fraction of completely solved programs. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SWE-bench MultilingualCoding89.1%
      Reported settings & source

      Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      SWE-bench MultimodalSupporting evidence54.7%
      Reported settings & source

      Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-12
      SWE-bench ProCoding81.2%Anthropic SWE-bench Pro reported system

      Effort: Not specified

      lab self-report

      Legacy self-report lacks a reviewed comparison configuration

      2026-09-01
      SWE-bench ProCoding81.2%
      Reported settings & source

      Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Terminal-Bench · 4.0Coding55.8%
      Reported settings & source

      Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

      Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12

      Published configuration

      Effort: Max

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Terminal-Bench 4.0 · 4.0Coding55.8%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench Science 0.1 · not specifiedHard reasoning52.6%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / Claude Fable 5.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench-Science · 0.1Coding52.6%
      Reported settings & source

      70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-12
      Toolathlon · Verified June2026Agentic77.8%
      Reported settings & source

      Pass@1;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      System card section8.15.5 confirms11of324trials were partly or fully completed by Opus4.8fallback; actual mixed-model execution, not merely available fallback.

      Reviewed 2026-09-12
      Toolathlon · Verified June2026Agentic81.5%
      Reported settings & source

      Pass@3;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      System card section8.15.5 confirms11of324trials were partly or fully completed by Opus4.8fallback; actual mixed-model execution, not merely available fallback.

      Reviewed 2026-09-12
      Toolathlon · Verified June2026Agentic73.1%
      Reported settings & source

      Pass³;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      System card section8.15.5 confirms11of324trials were partly or fully completed by Opus4.8fallback; actual mixed-model execution, not merely available fallback.

      Reviewed 2026-09-12
      Toolathlon · Verified June2026Agentic23.7 turns
      Reported settings & source

      average turns;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

      Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

      Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

      Published configuration

      Effort: Max

      lab self-report

      System card section8.15.5 confirms11of324trials were partly or fully completed by Opus4.8fallback; actual mixed-model execution, not merely available fallback.

      Reviewed 2026-09-12
      GDPval-AAAgentic
      LiveCodeBenchCoding
      LMArena Text ArenaHuman pref
      MMLU-ProKnowledge
      OSWorld-VerifiedAgentic
      SWE-bench VerifiedAgentic
      Terminal-Bench 2.1Agentic