RankingClaude Fable 5.1
Claude Fable 5.1
40 published benchmark measures · 16 benchmark families contribute across 5 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 5 effort levels across 63 benchmark/harness combinations →
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Auto thinking: 2 observations
- High: 1 observations
- Max: 25 observations
- Medium: 3 observations
- XHigh: 4 observations
- Not specified: 33 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
16 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Fable
- Released
- —
- Context
- 1,000,000 tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- Default Capability family coverage
/badge/claude-fable-5-1.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase | Supporting evidence | 1694Reported settings & sourceArtificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 1686Reported settings & sourceArtificial Analysis; xhigh effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 1611Reported settings & sourceArtificial Analysis; high effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 61.5%Reported settings & sourceArtificial Analysis; max effort; component rubric pass. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 2025 ratingReported settings & sourceArtificial Analysis; max effort; component analytical quality. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 1495 ratingReported settings & sourceArtificial Analysis; max effort; component presentation. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI · 1 | Hard reasoning | 97.5%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI · 2 | Hard reasoning | 90%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-1 · 1 | Hard reasoning | 97.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 90% | ARC Prize verifiedcontributes to capability | official board | 2026-09-01 |
| ARC-AGI-2 · 2 | Hard reasoning | 90%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1.1 · v4.1.1 | Supporting evidence | 65.7 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AutomationBench | Agentic | 31.4%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench · not specified | Agentic | 31.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 84.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled; three Anthropic evaluation modifications Provider-published result; comparator measurements are not automatically independently reproduced. Astra launch footnote5: Claude scores use three modifications described in the Fable5.1 system card; not same controlled setting. GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.437 voxel IoUReported settings & sourceRandom1000 of17900files; five runs; adaptive thinking max; no tools; corrected camera prompt, raw shapes accepted, last code fence parsed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.843 voxel IoUReported settings & sourceRandom1000 of17900files; five runs; adaptive thinking max; with tools; corrected camera prompt, raw shapes accepted, last code fence parsed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Chartography | Multimodal | 42.6%Reported settings & source100tasks; adaptive thinking max; five runs; no tools; tools condition has container, standard libraries and crop tool. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Chartography | Multimodal | 86.2%Reported settings & source100tasks; adaptive thinking max; five runs; with tools; tools condition has container, standard libraries and crop tool. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CursorBench · 3.2.0 | Supporting evidence | 73.4%Reported settings & sourceCursor production agent harness; independently measured by Cursor; max effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CursorBench · 3.2.0 | Supporting evidence | 68%Reported settings & sourceCursor production agent harness; medium effort; $3.53/task. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSWE · 1.1 | Coding | 67.4%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. 113 tasks; original hidden-test grading. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.3 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 67.4% | Anthropic DeepSWE1.1 reported system | lab self-report | 2026-09-01 |
| DeepSWE v1.1 · v1.1 | Coding | 67.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierCode · 1.1 Extended | Supporting evidence | 63.6%Reported settings & sourceCognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierCode · 1.1 Main | Supporting evidence | 50.9%Reported settings & sourceCognition agentic coding; composite functional and code-quality score. Fable5.1 medium effort; Fable5 xhigh. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierCode 1.1 Extended (score) · 1.1 | Supporting evidence | 63.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Main (score) · 1.1 | Supporting evidence | 50.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 87.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierSWE · 2 | Supporting evidence | 0.57 fractionReported settings & sourceProximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1853Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1835Reported settings & sourceArtificial Analysis; xhigh effort; same release board snapshot. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.3 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GMMLU | Knowledge | 94%Reported settings & sourceMean accuracy42languages; adaptive max; one trial; no tools/custom system prompts. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond | Hard reasoning | 93.737% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 93.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HealthBench | Supporting evidence | 66.7%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench | Supporting evidence | 60%Reported settings & sourceLength-adjusted score using GPT5.5card method; otherwise raw evaluation configuration. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional | Supporting evidence | 74.2%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional | Supporting evidence | 62.1%Reported settings & sourceLength-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional (length-adjusted) · not specified | Supporting evidence | 58.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted; OpenAI reproduction; GPT-5.4 grader; unclipped; Opus5 fallback for provider refusals Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Humanity's Last Exam | Hard reasoning | 60.9% | HLE no tools | lab self-report | 2026-09-01 |
| Humanity's Last Exam | Hard reasoning | 65% | HLE with tools | lab self-report | 2026-09-01 |
| Humanity’s Last Exam | Hard reasoning | 60.9%Reported settings & sourceFull2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Humanity’s Last Exam | Hard reasoning | 65%Reported settings & sourceFull2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Humanity's Last Exam (w/ tools) · not specified | Hard reasoning | 65%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Internal Database Migration Tasks · not specified | Supporting evidence | 57.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / Claude Fable 5.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Legal Agent Benchmark · 120-task held-out subset | Supporting evidence | 16.7%Reported settings & sourceall-pass; Artificial Analysis harness; xhigh effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Legal Agent Benchmark · 120-task held-out subset | Supporting evidence | 93.3%Reported settings & sourcecriterion-pass; Artificial Analysis harness; xhigh effort. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Legal Agent Benchmark · 1235-task public subset | Supporting evidence | 19.09%Reported settings & sourceall-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Legal Agent Benchmark · 1235-task public subset | Supporting evidence | 90.81%Reported settings & sourcecriterion-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MILU | Knowledge | 93%Reported settings & sourceMean accuracy11languages; adaptive max; five trials; no tools/custom system prompts. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OfficeQA | Agentic | 80.2%Reported settings & sourceExtracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OfficeQA Pro | Agentic | 69%Reported settings & sourceExtracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld · 2.0 August2026 task release | Agentic | 77.9%Reported settings & sourcepartial pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero. Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld · 2.0 August2026 task release | Agentic | 41.7%Reported settings & sourcestrict pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero. Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ProgramBench · 166 golden-task subset | Supporting evidence | 87.6%Reported settings & sourcemini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext. Hidden-test pass rate, not fraction of completely solved programs. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Multilingual | Coding | 89.1%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Multimodal | Supporting evidence | 54.7%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro | Coding | 81.2% | Anthropic SWE-bench Pro reported system | lab self-report | 2026-09-01 |
| SWE-bench Pro | Coding | 81.2%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench · 4.0 | Coding | 55.8%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 4.0 · 4.0 | Coding | 55.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Science 0.1 · not specified | Hard reasoning | 52.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / Claude Fable 5.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench-Science · 0.1 | Coding | 52.6%Reported settings & source70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 77.8%Reported settings & sourcePass@1;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 81.5%Reported settings & sourcePass@3;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 73.1%Reported settings & sourcePass³;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 23.7 turnsReported settings & sourceaverage turns;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA | Agentic | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |