RankingHumanity’s Last Exam

Humanity’s Last Exam

Data updated 24 Sept 2026

Bucket
Hard reasoning
Unit
percent
Direction
Higher is better
Version
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–8 of 8 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic60.9%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic57.8%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic56.6%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5Anthropic63.8%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Fable 5.1Anthropic65%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5Anthropic63.6%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-24

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-24
Claude Opus 5.5Anthropic64.4%
Reported settings & source

Full 2,500 questions; no tools; section 8.11.1 sets thinking to auto, a 1M total token cap, no compaction, and an Opus 4.6 grader.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's default note says adaptive max effort unless otherwise noted. Section 8.11.1 states thinking was set to auto for these runs.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Opus 5.5Anthropic67.7%
Reported settings & source

Full 2,500 questions; web search, web fetch, programmatic tool calling, and code execution; thinking set to auto; 1M total token cap; no compaction; Opus 4.6 grader; HLE source blocklist and contamination review.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 67.7% with tools.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22