RankingHumanity’s Last Exam

Humanity’s Last Exam

Data updated 29 Sept 2026

Bucket
Hard reasoning
Unit
percent
Direction
Higher is better
Version
—
Display harness
Claude Fable 5.1 and Claude Mythos 5.1 System Card
Board
https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–8 of 8 entries

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic60.9%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5Anthropic57.8%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5Anthropic56.6%
Reported settings & source

Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5Anthropic63.8%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Fable 5.1Anthropic65%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Fable is the safeguarded deployed configuration; selected tasks may use disclosed Opus fallback, except evaluations explicitly counting safety blocks as failures.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5Anthropic63.6%
Reported settings & source

Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-29

Claude Fable 5.1 and Claude Mythos 5.1 System Cardlab self-report2026-09-29
Claude Opus 5.5Anthropic64.4%
Reported settings & source

Full 2,500 questions; no tools; section 8.11.1 sets thinking to auto, a 1M total token cap, no compaction, and an Opus 4.6 grader.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Table 8.1.A's default note says adaptive max effort unless otherwise noted. Section 8.11.1 states thinking was set to auto for these runs.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22
Claude Opus 5.5Anthropic67.7%
Reported settings & source

Full 2,500 questions; web search, web fetch, programmatic tool calling, and code execution; thinking set to auto; 1M total token cap; no compaction; Opus 4.6 grader; HLE source blocklist and contamination review.

Provider-published lab self-report from the Claude Opus 5.5 system card (SHA256 7311c9c6bbb16d012f1c12c7418b05949fcf7ae3e30d2c40f22050074b2a7378). Not an independent board. Launch grid shows 67.7% with tools.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Claude Opus 5.5 System Card · Table 8.1.A; section 8.11.1 · reviewed 2026-09-22

Claude Opus 5.5 System Cardlab self-report2026-09-22