Compare published benchmarks

Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.

These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.

Category weights

Weights are relative. Categories share the total; each benchmark family shares its category equally. Related versions and settings divide their family’s allocation.

DeepSeek V4-Pro 0813 and Gemini 3.1 Pro

No unambiguous shared results under these settings. Inspect the available results below or choose another source.

Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.

BenchmarkDeepSeek V4-Pro 0813Gemini 3.1 ProComparison
HLE (wo / w tools) · source release snapshot; version not specified
Hard reasoning · Higher is better

42.7%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

60%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

No resultMissing result
Terminal Bench 2.1 · 2.1
Coding · Higher is better

87.9%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

No resultMissing result
NL2Repo · source release snapshot; version not specified
Coding · Higher is better

61.5%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: NL2Repo · reviewed 2026-09-12

No resultMissing result
Cybergym · source release snapshot; version not specified
Coding · Higher is better

83.3%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

No resultMissing result
DeepSWE · source release snapshot; version not specified
Coding · Higher is better

62.7%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

No resultMissing result
Toolathlon-Verified · source release snapshot; version not specified
Agentic · Higher is better

74.1%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

No resultMissing result
Agents' Last Exam · source release snapshot; version not specified
Agentic · Higher is better

25.7%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12

No resultMissing result
AutomationBench (Public) · source release snapshot; version not specified
Agentic · Higher is better

31.8%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

No resultMissing result
DSBench-FullStack † · source release snapshot; version not specified
Agentic · Higher is better

71.1%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result. † source footnote applies.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

No resultMissing result
DSBench-Hard † · source release snapshot; version not specified
Agentic · Higher is better

67.2%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result. † source footnote applies.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

No resultMissing result

Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations