The leaders, at a glance

UnifyBench ranking

Compare AI models and measured effort settings across published benchmarks.

Core work leaders

Experimental comparison score · 0–100

For the selected profile and sources. Scores are experimental; small gaps may not be meaningful. How scoring works · Effort evidence collected

Explore all 109 sources Independent evaluations and lab reports

Distinct source pages with published results in the catalog. Multiple results and page sections count as one source.

Every ranked row is a measured configuration. The default view shows the highest-scoring supported configuration per model. In Value, the budget is applied first. Select All efforts for the full comparison. How effort ranking works.

Customize ranking
Adjust weights

Full ranking

451 entries

Custom capability ranking of measured model configurations. Scores are experimental; missing results remain unknown.
RankModelScoreCapability coverageanalyst-agentbriefcaseaimeapex-agentsaa-lcrautomationbenchcritptenterprise-opsgdp-pdfgdpvalgmmlugpqaharveyhleifbenchitbenchlivecodebenchmath500mmlu-prommmu-proomnisciencescicodetau-benchterminal-bencharc-agideepsweswe-bench-proarc-agi-3frontiercodeenigma-evalswe-atlas-qnaswe-atlas-refactoringswe-atlas-test-writinglivebench-mathlivebench-codinglivebench-languagelivebench-datamulti-swe-benchlivebench-reasoninglivebench-instructionsswe-rebenchepoch-game-puzzlesfrontiermathsimpleqavals-finance-agentvals-code-migrationvals-legal-researchvals-excel-modelingvibe-codechartographyosworldcharxivswe-bench-multilingualdeepsearchqaofficeqamcp-atlastoolathlonmrcrbrowsecomp
Command A 03 2025 · Effort unspecifiedCohereNot scoredProvisional · 0%1 published · 0 families · 0 selected capabilities
Command A Reasoning 08 2025 · Effort unspecifiedCohereNot scoredProvisional · 0%10 published · 0 families · 0 selected capabilities
Command A Translate 08 2025 · Effort unspecifiedCohereNot scoredNo evidence0 published · 0 families · 0 selected capabilities
Command A Vision 07 2025 · Effort unspecifiedCohereNot scoredProvisional · 0%4 published · 0 families · 0 selected capabilities
Command Light · Effort unspecifiedCohereNot scoredNo evidence0 published · 0 families · 0 selected capabilities
Command R 03 2024 · Non-reasoningCohereNot scoredProvisional · 0%10 published · 0 families · 0 selected capabilities
Command R 08 2024 · Effort unspecifiedCohereNot scoredProvisional · 0%1 published · 0 families · 0 selected capabilities
Command R Plus 04 2024 · Effort unspecifiedCohereNot scoredProvisional · 0%1 published · 0 families · 0 selected capabilities
Command R Plus 08 2024 · Effort unspecifiedCohereNot scoredProvisional · 0%1 published · 0 families · 0 selected capabilities
Command R7B 12 2024 · Effort unspecifiedCohereNot scoredNo evidence0 published · 0 families · 0 selected capabilities
DeepSeek-R1-Distill-Qwen-7B · Effort unspecifiedDeepSeekNot scoredNo evidence0 published · 0 families · 0 selected capabilities
DeepSeek-V4-Flash-Vision-Exp · Effort unspecifiedDeepSeekNot scoredNo evidence0 published · 0 families · 0 selected capabilities
Devstral Medium 1.0 · Effort unspecifiedMistralNot scoredNo evidence0 published · 0 families · 0 selected capabilities
Devstral Small 1.1 · Effort unspecifiedMistralNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-0.3B-PT · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-21B-A3B-PT · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-21B-A3B-Thinking · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-turbo-128k · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-turbo-20260402 · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-turbo-32k · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-turbo-vl · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-turbo-vl-32k · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-VL-28B-A3B-PT · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-VL-28B-A3B-Thinking · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities
ERNIE-4.5-VL-424B-A47B-PT · Effort unspecifiedBaiduNot scoredNo evidence0 published · 0 families · 0 selected capabilities

Behind the ranking

273 entries with published results · 112 ranked · 0 provisional estimates.

Explore seven capabilities, the original results, and gaps in the evidence.

Cite this ranking

Include the model and effort setting, the selected profile, and the dated data downloads. A shared link uses the latest data; it does not freeze a ranking in time.

UnifyBench ranking
Data updated: 12 Sept 2026
Mode: custom; evidence: documented; configurations: best
Weights: v3;agentic:34,hard_reasoning:33,coding:33,human_pref:0,knowledge:0,multimodal:0,long_context:0
https://unifybench.ai/?mode=custom&page=11&w=v3

Catalog data · Effort measurements · Configuration scoring rules · Methodology

Catalog dated 12 Sept 2026. Effort evidence collected 12 Sept 2026. Collection dates are not evaluation dates.

What does the score mean?

The adjusted comparison score estimates performance against the same reference panel from matched published comparisons. Capability pools matched comparisons across families and checks aggregate breadth and independent corroboration. Custom profiles require support in every selected capability. Missing results never count as losses.

Family details show observed win shares against available reference peers. Those descriptive values are supporting evidence, not adjusted capability scores. Provider reports and independent results remain labeled. Matching reported settings do not establish equal inference effort, and score differences do not establish statistical significance.

Read the full methodology →