Compare published benchmarks
Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.
These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.
Kimi K3 and GPT-5.6 Sol
No unambiguous shared results under these settings. Inspect the available results below or choose another source.
Missing results never count as losses. Changing the source, selected models or weights can change the comparison. This is not a global model rank.
| Benchmark | Kimi K3 | GPT-5.6 Sol | Comparison |
|---|---|---|---|
| HLE (wo / w tools) · source release snapshot; version not specified | 43.5% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 56% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 | No result | Missing result |
| Terminal Bench 2.1 · 2.1 | 88.3% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | No result | Missing result |
| Cybergym · source release snapshot; version not specified | 80% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12 | No result | Missing result |
| DeepSWE · source release snapshot; version not specified | 67.5% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12 | No result | Missing result |
| Toolathlon-Verified · source release snapshot; version not specified | 76.5% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | No result | Missing result |
| Agents' Last Exam · source release snapshot; version not specified | 27.6% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | No result | Missing result |
| AutomationBench (Public) · source release snapshot; version not specified | 30.8% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12 | No result | Missing result |
| DSBench-FullStack † · source release snapshot; version not specified | 73.7% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12 | No result | Missing result |
| DSBench-Hard † · source release snapshot; version not specified | 63% DeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12 | No result | Missing result |
Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations