RankingLegal Agent Benchmark · 1235-task public subset

Legal Agent Benchmark · 1235-task public subset

Data updated 12 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
1235-task public subset
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic19.09%
Reported settings & source

all-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12
Claude Fable 5.1Anthropic90.81%
Reported settings & source

criterion-pass; five runs; adaptive max; internal bash/Python harness, Sonnet4.6judge;16defective tasks excluded; production safeguards/fallback.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.2 · reviewed 2026-09-12

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-12