Weighted consensus · Public benchmarks
LLM reasoning leaderboard
Compare AI reasoning models and measured effort settings across published reasoning benchmarks.
Reasoning leaders
Consensus across available benchmarks. Models cover different task areas; use a task ranking for a focused comparison.
No configuration yet meets the ranking requirements for this profile. Existing results and capability profiles remain available. Explore the evidence →
No supported leaders match these filters. Clear the search or explore the results below.
Experimental scores for the selected profile and sources. Small gaps may not be meaningful. How scoring works · Effort evidence collected
Explore all 119 source pages Independent evaluations and lab reports
Each model is counted once across effort levels. Configurations count documented model + effort settings, including unranked.
Original reports and benchmark leaderboards. Distinct source pages with published results in the catalog. Multiple results and page sections count as one source.
- ai.meta.com
- Amazon Nova 2 technical report
- anthropic.com
- ARC Prize: ARC-AGI-2
- ARC Prize: ARC-AGI-3 · provider-adapter
- arcprize.org
- Artificial Analysis: AA-AnalystAgent
- Artificial Analysis: AA-Briefcase Elo
- Artificial Analysis: AA-LCR v1.1
- Artificial Analysis: AA-Omniscience Index
- Artificial Analysis: AIME 2025
- Artificial Analysis: APEX-Agents-AA
- Artificial Analysis: CritPt
- Artificial Analysis: EnterpriseOps-Gym-AA
- Artificial Analysis: GDP.pdf: All-pass
- Artificial Analysis: Global-MMLU-Lite
- Artificial Analysis: Harvey LAB-AA: Criterion Pass Rate
- Artificial Analysis: Humanity's Last Exam
- Artificial Analysis: IFBench
- Artificial Analysis: ITBench-AA
- Artificial Analysis: LiveCodeBench
- Artificial Analysis: MATH-500
- Artificial Analysis: MLCR-AA
- Artificial Analysis: MMLU-Pro
- Artificial Analysis: MMMU-Pro
- Artificial Analysis: SciCode
- Artificial Analysis: Terminal-Bench Hard
- Artificial Analysis: Terminal-Bench v2.1
- Artificial Analysis: Terminal-Bench v4.0
- Artificial Analysis: Terminal-Bench-Science 0.1
- Artificial Analysis: 𝜏²-Bench Telecom
- Artificial Analysis: 𝜏³-Banking
- artificialanalysis.ai
- artificialanalysis.ai
- artificialanalysis.ai
- artificialanalysis.ai
- artificialanalysis.ai
- artificialanalysis.ai
- artificialanalysis.ai
- benchmarks.bespokelabs.ai
- bughunt.productcompass.pm
- Claude Fable 5.1 and Claude Mythos 5.1 System Card
- Claude Opus 5.5 System Card
- Cognition: FrontierCode 1.1 · extended-chisel
- Command A+ launch benchmarks
- Datacurve: DeepSWE v1.1
- deepmind.google
- deepseek-ai/DeepSeek-V4-Pro-0813
- deepswe.datacurve.ai
- Epoch AI: Epoch · Chess Puzzles · v1.0.0
- Epoch AI: Epoch · FrontierMath Tier 4 · v2.0.0
- Epoch AI: Epoch · FrontierMath Tiers 1–3 · v2.0.0
- Epoch AI: Epoch · GPQA diamond · v1.0.0
- Epoch AI: Epoch · Mystery Game Puzzles · v1.0.4 · Tool submission
- Epoch AI: Epoch · OTIS Mock AIME 2024-2025 · v1.0.0
- Epoch AI: Epoch · SimpleQA Verified · v1.0.0
- Gemini 3.8 Flash launch performance
- Gemini3.8Flash Model Card
- GPT-5.6: Frontier intelligence that scales with your ambition
- GPT-6 Astra System Card
- GPT-6 Astra: A new generation of intelligence
- Grok 4.6 launch evaluations
- Grok 4.7 launch comparison
- huggingface.co
- huggingface.co
- huggingface.co
- Introducing GPT-6 Sol and Luna
- Introducing GPT-6.1 Sol
- Kimi K2.6 official model card
- kimi.ai
- LiveBench: LiveBench · AMPS Hard · 2026-06-25
- LiveCodeBench: LiveCodeBench
- livecodebench.github.io
- lmarena.ai
- MAI-Thinking-1 technical report
- MiniMax M3 model card benchmark figure
- Mistral Medium 3.5 model card performance charts
- moonshotai/Kimi-K3
- Muse Spark 1.1 evaluation report Figure44
- Muse Spark1.3 evaluation methodology
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- os-world.github.io
- qwen.ai
- Qwen/Qwen3.8-2.4T-A95B
- realswe.withspecific.com
- Scale: EnigmaEval · Scale official evaluation
- Scale: HLE · full · Scale official evaluation
- Scale: HLE · text only · Scale official evaluation
- Scale: SWE Atlas · QnA · claude code
- Scale: SWE Atlas · Refactoring · claude code
- Scale: SWE Atlas · Test Writing · claude code
- Scale: SWE-bench Pro public
- scale.com
- scale.com
- SWE-rebench: SWE-rebench · 2025-08 · tools
- swebench.com
- tbench.ai
- tbench.ai
- Terminal-Bench: Terminal-Bench 4.0 · claudecode
- Vals AI: Vals AI · Code Migration v1
- Vals AI: Vals AI · Excel Modeling Benchmark v1
- Vals AI: Vals AI · Finance Agent v2
- Vals AI: Vals AI · Harvey Legal Agent Benchmark v1
- Vals AI: Vals AI · Legal Research Bench v1
- Vals AI: Vals AI · MMLU-Pro v1
- Vals AI: Vals AI · Tax Agent Bench v1
- Vals AI: Vals AI · Terminal-Bench 2.1
- Vals AI: Vals AI · Vibe Code Bench v1.1
- vals.ai
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- vulcanbench.com
- x.com
- zai-org/GLM-5.3
One entry per model, using its highest-scoring eligible effort setting. How effort ranking works.
Customize ranking Sources, profiles and evidence details
Full ranking
222 entriesNo capabilities are weighted. Increase a weight to compare model performance.
| Rank | Model | Score | Capability coverage |
|---|---|---|---|
| 51 | Qwen3.8-27B · XHighQwen | 33.3 | Eligible · 100%3 families · 1 selected capabilities |
| 52 | DeepSeek V4 Flash 0424 · MaxDeepSeek | 32.6 | Eligible · 100%3 families · 1 selected capabilities |
| 53 | Claude Sonnet 4.6 · HighAnthropic | 32.5 | Eligible · 100%4 families · 1 selected capabilities |
| 54 | Hy3 · ReasoningTencent | 31.9 | Eligible · 100%3 families · 1 selected capabilities |
| 55 | GPT-6 Luna · MaxOpenAI | 31.1 | Eligible · 100%4 families · 1 selected capabilities |
| 56 | Solar Pro 4 · ReasoningUpstage | 30.2 | Eligible · 100%3 families · 1 selected capabilities |
| 57 | DeepSeek-V3.2-Speciale · ReasoningDeepSeek | 29.3 | Eligible · 100%4 families · 1 selected capabilities |
| 58 | MiMo-V2.5-Pro · ReasoningXiaomi | 28.9 | Eligible · 100%3 families · 1 selected capabilities |
| 59 | GLM-5.1 · ReasoningZ.ai | 28.3 | Eligible · 100%3 families · 1 selected capabilities |
| 60 | Solar Open2 250B · ReasoningUpstage | 27.4 | Eligible · 100%3 families · 1 selected capabilities |
| 61 | Qwen3.5-397B-A17B · ReasoningQwen | 26.4 | Eligible · 100%3 families · 1 selected capabilities |
| 62 | Hy3-preview · ReasoningTencent | 26.1 | Eligible · 100%3 families · 1 selected capabilities |
| 63 | Qwen3.6-plus · ReasoningQwen | 25.8 | Eligible · 100%3 families · 1 selected capabilities |
| 64 | NVIDIA Nemotron 3 Ultra · ReasoningNVIDIA | 25.6 | Eligible · 100%3 families · 1 selected capabilities |
| 65 | GPT-5.1 · HighOpenAI | 24.1 | Eligible · 100%5 families · 1 selected capabilities |
| 66 | MiMo-V2.5 · ReasoningXiaomi | 24.0 | Eligible · 100%3 families · 1 selected capabilities |
No models match these filters. Clear search and coverage above, or choose All models and evidence to include unsupported entries.
About this ranking
This profile gives hard reasoning all the weight. It compares identified configurations using their own measured results, rather than borrowing results from another effort setting. Rankings require support in this capability. Matching effort labels across providers do not imply equal compute budgets.
View the Capability ranking →Behind the ranking
279 entries with published results · 222 ranked · 0 provisional estimates.
Explore seven capabilities, the original results, and gaps in the evidence.
Cite this ranking
Include the model and effort setting, the selected profile, and the dated data downloads. A shared link uses the latest data; it does not freeze a ranking in time.
UnifyBench ranking Data updated: 30 Sept 2026 Mode: custom; evidence: documented; configurations: best Weights: v3;agentic:0,hard_reasoning:100,coding:0,human_pref:0,knowledge:0,multimodal:0,long_context:0 https://unifybench.ai/rankings/reasoning?evidence=documented&page=3&view=shortlist
Catalog data · Effort measurements · Configuration scoring rules · Methodology
What does the score mean?
The adjusted comparison score estimates performance against the same reference panel from matched published comparisons. Capability pools matched comparisons across families and checks aggregate breadth and independent source coverage. Custom profiles require support in every selected capability. Missing results never count as losses.
Family details show observed win shares against available reference peers. Those descriptive values are supporting evidence, not adjusted capability scores. Provider reports and independent results remain labeled. Matching reported settings do not establish equal inference effort, and score differences do not establish statistical significance.
Read the full methodology →