RankingDeepSeek V4-Pro 0813

DeepSeek V4-Pro 0813

Data updated 12 Sept 2026

27 published benchmark measures · 9 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
DeepSeek V4-Pro 0813 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100DeepSeek V4-Pro 0813 · Agentic: 65.1 · SupportedDeepSeek V4-Pro 0813 · Hard reasoning: 53.9 · SupportedDeepSeek V4-Pro 0813 · Coding: 61.0 · SupportedDeepSeek V4-Pro 0813 · Human pref: 31.4 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic3 families · 1 with independent evidence · Supported65.1

3 core families; 44 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 47.1–82.7; 0/11 scenarios unsupported. Without one publisher: 57.3–71.3; 0/15 unsupported. Smoothing check: 61.7–66.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning3 families · 2 with independent evidence · Supported53.9

3 core families; 59 direct opponents across 12 labs.

Disputed order: matched results across at least two families give the opposite order against 1 peers. Different test mixes and indirect comparisons can cause this. Matched comparison 1

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 48.6–55.8; 0/7 scenarios unsupported. Without one publisher: 47.4–59.1; 0/16 unsupported. Smoothing check: 53.0–54.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 1 with independent evidence · Supported61.0

2 core families; 27 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 55.6–62.2; 2/7 scenarios unsupported. Without one publisher: 54.1–65.1; 0/20 unsupported. Smoothing check: 59.0–61.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported31.4

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 22.0–40.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 81.6 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    MultimodalNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Long contextNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

        Compare 5 effort levels across 37 benchmark/harness combinations →

        Reported effort · Max + unspecified

        Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

        Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

        Inspect each result and its source ↓ · Download effort evidence

        Compare capability profiles →

        Score contributions and missing evidence

        9 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.

        Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

        Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

        Model information & shareable badge
        Lab
        DeepSeek
        Catalog status
        api-listed
        Availability
        Documented provider API and open weights for self-hosting
        Family
        DeepSeek V4
        Released
        Context
        1,000,000 tokens
        License
        mit
        Model card
        https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
        Default Capability family coverage
        Documented-evidence family coverage/badge/deepseek-v4-pro.svg
        Benchmark scores & sources

        Original results, evaluation harnesses, and evidence behind this model.

        BenchmarkBucketScoreHarnessEvidenceSource-recorded date
        Agents' Last Exam · source release snapshot; version not specifiedSupporting evidence25.7%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence25.7%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        ARC-AGI-2Hard reasoning61.3%ARC Prize verified

        Effort: Not specified

        contributes to capability
        official board2026-08-13
        Artificial Analysis Intelligence IndexSupporting evidence53 AA Intelligence Index

        Effort: Not specified

        official board

        Supporting evidence outside the reviewed capability core

        2026-09-02
        AutomationBench (Public) · source release snapshot; version not specifiedAgentic31.8%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        AutomationBench (v1.0.6) · v1.0.6Agentic43.2%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        BrowseCompAgentic83.4%BrowseComp reported

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        Cybergym · source release snapshot; version not specifiedSupporting evidence83.3%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

        Published configuration

        Effort: Max

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        CyberGym · source release snapshot; version not specifiedSupporting evidence83.3%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        DeepSWE · source release snapshot; version not specifiedCoding62.7%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-12
        DeepSWE (v1.1) · v1.1Coding62.7%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        DeepSWE v1.1Coding62.8%DeepSWE v1.1 reported

        Effort: Not specified

        contributes to capability
        official board2026-09-03
        DSBench-FullStack † · source release snapshot; version not specifiedSupporting evidence71.1%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result. † source footnote applies.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        DSBench-Hard † · source release snapshot; version not specifiedSupporting evidence67.2%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result. † source footnote applies.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        GDPval-AAAgentic1493Artificial Analysis GDPval-AA

        Effort: Not specified

        contributes to capability
        official board2026-09-12
        GDPval-AA v2 · source release snapshot; version not specifiedAgentic1590
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        Externally evaluated result, attributed in the source footnote.

        Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

        zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

        Reviewed 2026-09-12
        GPQA DiamondHard reasoning92.828%GPQA Diamond reported

        Effort: Not specified

        contributes to capability
        independent repro2026-09-11
        HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning42.7%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        HLE (wo / w tools) · source release snapshot; version not specifiedHard reasoning60%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        HLE w/ Tools · source release snapshot; version not specifiedHard reasoning60%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        Humanity's Last ExamHard reasoning42.7%HLE no tools

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        LiveCodeBenchCoding93.5%LiveCodeBench pass@1

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        LMArena Text ArenaHuman pref1457LMArena Text

        Effort: Not specified

        contributes to capability
        official board2026-09-11
        NL2Repo · source release snapshot; version not specifiedSupporting evidence61.5%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: NL2Repo · reviewed 2026-09-12

        Published configuration

        Effort: Max

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        NL2Repo · source release snapshot; version not specifiedSupporting evidence61.1%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        SWE-bench ProCoding55.4%SWE-bench Pro reported

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        SWE-bench VerifiedCoding80.6%SWE-bench Verified official

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        Terminal Bench 2.1 · 2.1Coding87.9%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-12
        Terminal Bench 2.1 · 2.1Coding87.9%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        Terminal-Bench 2.1Coding87.9%Terminal-Bench 2.1 reported

        Effort: Not specified

        lab self-report

        Legacy self-report lacks a reviewed comparison configuration

        2026-08-13
        Toolathlon Verified · source release snapshot; version not specifiedAgentic74.1%
        Reported settings & source

        GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

        First-party reported result.

        zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        Toolathlon-Verified · source release snapshot; version not specifiedAgentic74.1%
        Reported settings & source

        DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

        First-party reported result.

        deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        MMLU-ProKnowledge
        OSWorld-VerifiedAgentic