The leaders, at a glance

Agentic AI model ranking

Compare AI models and effort settings for agentic work using sourced task evaluations and visible evidence coverage.

Agentic leaders

Experimental comparison score · 0–100

For the selected profile and sources. Scores are experimental; small gaps may not be meaningful. How scoring works · Effort evidence collected

Explore all 109 sources Independent evaluations and lab reports

Distinct source pages with published results in the catalog. Multiple results and page sections count as one source.

Every ranked row is a measured configuration. The default view shows the highest-scoring supported configuration per model. In Value, the budget is applied first. Select All efforts for the full comparison. How effort ranking works.

Customize ranking
Adjust weights

Full ranking

127 entries

Custom capability ranking of measured model configurations. Scores are experimental; missing results remain unknown.
RankModelScoreCapability coverageanalyst-agentbriefcaseaimeapex-agentsaa-lcrautomationbenchcritptenterprise-opsgdp-pdfgdpvalgmmlugpqaharveyhleifbenchitbenchlivecodebenchmath500mmlu-prommmu-proomnisciencescicodetau-benchterminal-bencharc-agideepsweswe-bench-proarc-agi-3frontiercodeenigma-evalswe-atlas-qnaswe-atlas-refactoringswe-atlas-test-writinglivebench-mathlivebench-codinglivebench-languagelivebench-datamulti-swe-benchlivebench-reasoninglivebench-instructionsswe-rebenchepoch-game-puzzlesfrontiermathsimpleqavals-finance-agentvals-code-migrationvals-legal-researchvals-excel-modelingvibe-codechartographyosworldcharxivswe-bench-multilingualdeepsearchqaofficeqamcp-atlastoolathlonmrcrbrowsecomp
126Gemma 3 27B IT · Non-reasoningGoogle1.2Eligible · 100%22 published · 4 families · 1 selected capabilities20.0108/108 peersIndependent evidence12.6197/197 peersIndependent evidence3.189/89 peersIndependent evidence25.2200/200 peersIndependent evidence5.488/88 peersIndependent evidence0.9123/123 peersIndependent evidence30.763/63 peersIndependent evidence12.5212/212 peersIndependent evidence18.7208/208 peersIndependent evidence19.2178/178 peersIndependent evidence8.1112/112 peersIndependent evidence45.160/60 peersIndependent evidence22.4114/114 peersIndependent evidence18.2114/114 peersIndependent evidence16.5198/198 peersIndependent evidence5.291/91 peersIndependent evidence5.4202/202 peersIndependent evidence20.3201/201 peersIndependent evidence
127Phi-4-mini-instruct · Non-reasoningMicrosoft0.6Eligible · 100%15 published · 2 families · 1 selected capabilities4.9108/108 peersIndependent evidence19.7197/197 peersIndependent evidence25.5200/200 peersIndependent evidence0.0123/123 peersIndependent evidence6.5212/212 peersIndependent evidence19.3208/208 peersIndependent evidence3.8178/178 peersIndependent evidence5.8112/112 peersIndependent evidence11.560/60 peersIndependent evidence6.9114/114 peersIndependent evidence9.0198/198 peersIndependent evidence6.4171/171 peersIndependent evidence3.4201/201 peersIndependent evidence

About this ranking

This profile gives agentic work all the weight. Published agent-task evaluations contribute through matched benchmark families. Harnesses and reported settings remain attached to the evidence. These results describe the evaluated configurations; they do not establish how a model will perform in every workflow.

View the Capability ranking →

Behind the ranking

273 entries with published results · 127 ranked · 0 provisional estimates.

Explore seven capabilities, the original results, and gaps in the evidence.

Cite this ranking

Include the model and effort setting, the selected profile, and the dated data downloads. A shared link uses the latest data; it does not freeze a ranking in time.

UnifyBench ranking
Data updated: 12 Sept 2026
Mode: custom; evidence: documented; configurations: best
Weights: v3;agentic:100,hard_reasoning:0,coding:0,human_pref:0,knowledge:0,multimodal:0,long_context:0
https://unifybench.ai/rankings/agentic?page=6

Catalog data · Effort measurements · Configuration scoring rules · Methodology

Catalog dated 12 Sept 2026. Effort evidence collected 12 Sept 2026. Collection dates are not evaluation dates.

What does the score mean?

The adjusted comparison score estimates performance against the same reference panel from matched published comparisons. Capability pools matched comparisons across families and checks aggregate breadth and independent corroboration. Custom profiles require support in every selected capability. Missing results never count as losses.

Family details show observed win shares against available reference peers. Those descriptive values are supporting evidence, not adjusted capability scores. Provider reports and independent results remain labeled. Matching reported settings do not establish equal inference effort, and score differences do not establish statistical significance.

Read the full methodology →