RankingHealthBench
HealthBench
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- —
- Display harness
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Board
- https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–6 of 6 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Opus 5Anthropic | 67.1%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5.1Anthropic | 66.7%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5Anthropic | 61.2%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Fable 5.1Anthropic | 60%Reported settings & sourceLength-adjusted score using GPT5.5card method; otherwise raw evaluation configuration. Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17.1 · reviewed 2026-09-24 | Claude Fable 5.1 and Claude Mythos 5.1 System Card | lab self-report | 2026-09-24 |
| Claude Opus 5.5Anthropic | 68.1%Reported settings & sourceRaw rubric score; adaptive thinking at max effort; five trials; no tools or custom system prompt; Opus 4.8 grader; safety classifiers with refusal fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22 | Claude Opus 5.5 System Card | lab self-report | 2026-09-22 |
| Claude Opus 5.5Anthropic | 60.6%Reported settings & sourceLength-adjusted score using the GPT-5.5 system-card method; otherwise the raw HealthBench configuration: adaptive max, five trials, no tools, Opus 4.8 grader, safety fallback to Opus 5. Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Claude Opus 5.5 System Card · section 8.15.1 · reviewed 2026-09-22 | Claude Opus 5.5 System Card | lab self-report | 2026-09-22 |