RankingBrowseComp · source release snapshot; version not specified
BrowseComp · source release snapshot; version not specified
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- moonshotai/Kimi-K3
- Board
- https://huggingface.co/moonshotai/Kimi-K3
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Kimi K3Moonshot | 91.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 90.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 88%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.5OpenAI | 84.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 84.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.7Anthropic | 79.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Claude Sonnet 4.5Anthropic | 43.9%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Claude Sonnet 4.6Anthropic | 74.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Claude Sonnet 4.6Anthropic | 74.7%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 85.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 79.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 79.3%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 59.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 59.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 84.4%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 83.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 61.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 61.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| MiniMax-M2.7MiniMax | 76.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| MiniMax-M2.7MiniMax | 54.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 54.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| MiniMax-M3MiniMax | 83.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | MiniMax M3 model card benchmark figure | lab self-report | 2026-09-06 |
| Mistral Medium 3.5Mistral | 48.6%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| NVIDIA Nemotron 3 UltraNVIDIA | 44.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 44.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| Qwen3.5-397B-A17BQwen | 78.6%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Qwen3.5-397B-A17BQwen | 40.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 40.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |