Compare published benchmarks

Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.

These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.

Category weights

Weights are relative. Categories share the total; each benchmark family shares its category equally. Related versions and settings divide their family’s allocation.

Kimi K3 and GPT-5.6 Sol

No unambiguous shared results under these settings. Inspect the available results below or choose another source.

Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.

BenchmarkKimi K3GPT-5.6 SolComparison
HLE (wo / w tools) · source release snapshot; version not specified
Hard reasoning · Higher is better

43.5%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

56%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12

No resultMissing result
Terminal Bench 2.1 · 2.1
Coding · Higher is better

88.3%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

No resultMissing result
Cybergym · source release snapshot; version not specified
Coding · Higher is better

80%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12

No resultMissing result
DeepSWE · source release snapshot; version not specified
Coding · Higher is better

67.5%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12

No resultMissing result
Toolathlon-Verified · source release snapshot; version not specified
Agentic · Higher is better

76.5%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12

No resultMissing result
Agents' Last Exam · source release snapshot; version not specified
Agentic · Higher is better

27.6%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12

No resultMissing result
AutomationBench (Public) · source release snapshot; version not specified
Agentic · Higher is better

30.8%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12

No resultMissing result
DSBench-FullStack † · source release snapshot; version not specified
Agentic · Higher is better

73.7%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result. † source footnote applies.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12

No resultMissing result
DSBench-Hard † · source release snapshot; version not specified
Agentic · Higher is better

63%

Reported settings & source

DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card.

First-party reported result. † source footnote applies.

deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12

No resultMissing result

Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations