RankingHumanity's Last Exam
Humanity's Last Exam
- Bucket
- Hard reasoning
- Unit
- percent
- Direction
- Higher is better
- Version
- full
- Display harness
- HLE no tools
- Board
- https://lastexam.ai/
Compare published benchmark results with category weights →
Models
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 63.8% | HLE with tools | lab self-report | 2026-09-01 |
| Claude Fable 5.1Anthropic | 65% | HLE with tools | lab self-report | 2026-09-01 |
| Claude Opus 5Anthropic | 63.6% | HLE with tools | lab self-report | 2026-09-01 |
| Claude Opus 5Anthropic | 64.7% | HLE with tools | lab self-report | 2026-07-24 |