RankingAutomationBench
AutomationBench
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- —
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
Compare published benchmark results with category weights →
Models
1–5 of 5 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 31.4%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Opus 5Anthropic | 26.9%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| GPT-5.6 SolOpenAI | 19.6%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5Anthropic | 17.05%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Opus 5.5Anthropic | 40%Reported settings & sourceZapier private held-out leaderboard; simulated business workflows; every deterministic assertion must pass; max effort. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 40.0%. Launch footnote 2 says Zapier ran these without fallback models and counted safeguard interventions as failures. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · Table 8.1.A; section 8.14.6 · reviewed 2026-09-22 | Claude Opus 5.5 System Card | lab self-report | 2026-09-22 |