Weighted consensus · Public benchmarks

LLM reasoning leaderboard

Compare AI reasoning models and measured effort settings across published reasoning benchmarks.

Reasoning leaders

Comparison score · 0–100

Experimental scores for the selected profile and sources. Small gaps may not be meaningful. How scoring works · Effort evidence collected

Explore all 119 source pages Independent evaluations and lab reports

Each model is counted once across effort levels. Configurations count documented model + effort settings, including unranked.

Original reports and benchmark leaderboards. Distinct source pages with published results in the catalog. Multiple results and page sections count as one source.

One entry per model, using its highest-scoring eligible effort setting. How effort ranking works.

Customize ranking Sources, profiles and evidence details

Evidence only MultimodalLong contextHuman prefAll capabilitiesThese profiles do not yet have enough comparable evidence for a ranking.

Adjust weights

Full ranking

222 entries

Custom capability ranking of measured model configurations. Scores are experimental; missing results remain unknown.
RankModelScoreCapability coverage
76Claude Sonnet 4.5 · Thinking 32KAnthropic19.8Eligible · 100%4 families · 1 selected capabilities
77Gemini 3.1 Flash-Lite · HighGoogle19.5Eligible · 100%4 families · 1 selected capabilities
78MiniMax-M2.5 · ReasoningMiniMax19.3Eligible · 100%3 families · 1 selected capabilities
79Qwen3.5-27B · ReasoningQwen19.2Eligible · 100%3 families · 1 selected capabilities
80Qwen3.5-122B-A10B · ReasoningQwen18.9Eligible · 100%3 families · 1 selected capabilities
81Qwen3.6-27B · ReasoningQwen18.8Eligible · 100%3 families · 1 selected capabilities
82o3 · ReasoningOpenAI18.5Eligible · 100%5 families · 1 selected capabilities
83NVIDIA-Nemotron-3-Super-120B-A12B-BF16 · ReasoningNVIDIA18.1Eligible · 100%3 families · 1 selected capabilities
84Step 3.7 Flash · ReasoningStepFun18.0Eligible · 100%3 families · 1 selected capabilities
85Qwen3.5-35B-A3B · ReasoningQwen17.9Eligible · 100%3 families · 1 selected capabilities
86GLM-5-Turbo · ReasoningZ.ai17.4Eligible · 100%3 families · 1 selected capabilities
87Gemini 2.5 Pro · ReasoningGoogle16.8Eligible · 100%5 families · 1 selected capabilities
88Gemini 3.5 Flash-Lite · HighGoogle16.8Eligible · 100%7 families · 1 selected capabilities
89Qwen3.6-flash · Non-reasoningQwen16.3Eligible · 100%2 families · 1 selected capabilities
90K EXAONE 2.0 750B A37B · ReasoningLG AI Research16.1Eligible · 100%3 families · 1 selected capabilities
91Qwen3.6-35B-A3B · ReasoningQwen15.2Eligible · 100%3 families · 1 selected capabilities

More models on this page

About this ranking

This profile gives hard reasoning all the weight. It compares identified configurations using their own measured results, rather than borrowing results from another effort setting. Rankings require support in this capability. Matching effort labels across providers do not imply equal compute budgets.

View the Capability ranking →

Behind the ranking

279 entries with published results · 222 ranked · 0 provisional estimates.

Explore seven capabilities, the original results, and gaps in the evidence.

Cite this ranking

Include the model and effort setting, the selected profile, and the dated data downloads. A shared link uses the latest data; it does not freeze a ranking in time.

UnifyBench ranking
Data updated: 30 Sept 2026
Mode: custom; evidence: documented; configurations: best
Weights: v3;agentic:0,hard_reasoning:100,coding:0,human_pref:0,knowledge:0,multimodal:0,long_context:0
https://unifybench.ai/rankings/reasoning?evidence=documented&page=4&view=shortlist

Catalog data · Effort measurements · Configuration scoring rules · Methodology

Catalog dated 30 Sept 2026. Effort evidence collected 30 Sept 2026. Collection dates are not evaluation dates.

What does the score mean?

The adjusted comparison score estimates performance against the same reference panel from matched published comparisons. Capability pools matched comparisons across families and checks aggregate breadth and independent source coverage. Custom profiles require support in every selected capability. Missing results never count as losses.

Family details show observed win shares against available reference peers. Those descriptive values are supporting evidence, not adjusted capability scores. Provider reports and independent results remain labeled. Matching reported settings do not establish equal inference effort, and score differences do not establish statistical significance.

Read the full methodology →