RankingClaude Opus 4.8
Claude Opus 4.8
108 published benchmark measures · 22 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 6 effort levels across 73 benchmark/harness combinations →
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 1 observations
- Max: 50 observations
- Not specified: 106 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
22 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Opus
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- Default Capability family coverage
/badge/claude-opus-4-8.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| $OneMillion-Bench (expert score) · source release snapshot; version not specified | Supporting evidence | 41.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase (Elo) · source release snapshot; version not specified | Supporting evidence | 1354Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-LCR · source release snapshot; version not specified | Long context | 67.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam · not specified | Supporting evidence | 45.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agents' Last Exam · source release snapshot; version not specified | Supporting evidence | 27%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam · source release snapshot; version not specified | Supporting evidence | 25.7%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (ALE-CLI) · source release snapshot; version not specified | Supporting evidence | 25.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 27%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 45.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AndroidBench · source release snapshot; version not specified | Supporting evidence | 69.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 39.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-2 | Hard reasoning | 72.1% | ARC Prize verifiedcontributes to capability | official board | 2026-06-01 |
| ARC-AGI-3 · 3 | Hard reasoning | 1.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; high reasoning, not max Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.1 · v1.1 | Supporting evidence | 72.5 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1 · v4.1 | Supporting evidence | 55.7 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Automation-Bench (Pass@1) · source release snapshot; version not specified | Agentic | 27.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench · not specified | Agentic | 15.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · source release snapshot; version not specified | Agentic | 27.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (Public) · source release snapshot; version not specified | Agentic | 27.2%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (v1.0.6) · v1.0.6 | Agentic | 41%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| BabyVision w/ python · source release snapshot; version not specified | Supporting evidence | 81.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BabyVision with tools · source release snapshot; version not specified | Supporting evidence | 81.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 27.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD (python tool) · not specified | Supporting evidence | 51.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Big Finance Bench · not specified | Supporting evidence | 44%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · not specified | Agentic | 84.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 84.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) · source release snapshot; version not specified | Multimodal | 80.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) · source release snapshot; version not specified | Multimodal | 89.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv Reasoning with tools · source release snapshot; version not specified | Multimodal | 89.9%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| CorpFin v2 · source release snapshot; version not specified | Supporting evidence | 66.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CoWorkBench · source release snapshot; version not specified | Supporting evidence | 72.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CritPt · source release snapshot; version not specified | Supporting evidence | 20.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Cybergym · source release snapshot; version not specified | Supporting evidence | 78.3%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Cybergym · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CyberGym · source release snapshot; version not specified | Supporting evidence | 78.1%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSearchQA · source release snapshot; version not specified | Agentic | 84.3%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA (F1) · source release snapshot; version not specified | Agentic | 93.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: DeepSearchQA (F1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE · source release snapshot; version not specified | Coding | 59%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE · source release snapshot; version not specified | Coding | 58%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE (v1.1) · v1.1 | Coding | 58%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE 1.1 · 1.1 | Coding | 59%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 59% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 59%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 · v1.1 | Coding | 59%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DSBench-FullStack † · source release snapshot; version not specified | Supporting evidence | 71.6%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DSBench-Hard † · source release snapshot; version not specified | Supporting evidence | 71.7%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitBench · not specified | Supporting evidence | 40%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · source release snapshot; version not specified | Supporting evidence | 40%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 80 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 120 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 53.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 53.9%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 1-3 (v2) · v2 | Hard reasoning | 80%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 56.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 70%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 66.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 66.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| gdp.pdf · not specified | Supporting evidence | 22.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 · source release snapshot; version not specified | Agentic | 1588Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 · v2 | Agentic | 1600Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 (Elo) · source release snapshot; version not specified | Agentic | 1593Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 Elo · source release snapshot; version not specified | Agentic | 1600Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · not specified | Supporting evidence | 16%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · not specified | Hard reasoning | 92%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 92%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 91%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GraphWalks BFS 1mil f1 · not specified | Long context | 68.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GraphWalks BFS 256k f1 · not specified | Long context | 85.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey Lab-AA · source release snapshot; version not specified | Supporting evidence | 91.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench · source release snapshot; version not specified | Supporting evidence | 52.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional · not specified | Supporting evidence | 53%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · source release snapshot; version not specified | Supporting evidence | 55.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE · source release snapshot; version not specified | Hard reasoning | 45.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE (wo / w tools) · source release snapshot; version not specified | Hard reasoning | 49.8%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE (wo / w tools) · source release snapshot; version not specified | Hard reasoning | 57.9%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ tools · source release snapshot; version not specified | Hard reasoning | 57.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ Tools · source release snapshot; version not specified | Hard reasoning | 57.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE with tools · source release snapshot; version not specified | Hard reasoning | 57.9%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE without tools · source release snapshot; version not specified | Hard reasoning | 49.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE-Full · source release snapshot; version not specified | Hard reasoning | 49.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE-Full · source release snapshot; version not specified | Hard reasoning | 57.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| IFBench · source release snapshot; version not specified | Supporting evidence | 62.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 48.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 48.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 48.4%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Kimi Code Bench 2.0 · 2.0 | Supporting evidence | 71.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Legal Research Bench · source release snapshot; version not specified | Supporting evidence | 43.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LifeSciBench · not specified | Supporting evidence | 53.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1473 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| LongBench v2 · source release snapshot; version not specified | Long context | 69.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Management Consulting Tasks (Internal) · not specified | Supporting evidence | 31.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / Claude Opus 4.8 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MathVision · source release snapshot; version not specified | Supporting evidence | 86.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MathVision · source release snapshot; version not specified | Supporting evidence | 97.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MCP-Atlas · source release snapshot; version not specified | Agentic | 83.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MCPAtlas · source release snapshot; version not specified | Agentic | 82.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MCPMark-Verified · source release snapshot; version not specified | Supporting evidence | 76.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 42.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 42.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 78.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 82.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMVU · source release snapshot; version not specified | Supporting evidence | 79.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MRCR v2 256K (8-needle) · source release snapshot; version not specified | Long context | 83.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 69.7%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: NL2Repo · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 69.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| NL2Repo-Bench · source release snapshot; version not specified | Supporting evidence | 69.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: NL2Repo-Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OfficeQA Pro · source release snapshot; version not specified | Agentic | 63.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OmniDocBench · source release snapshot; version not specified | Supporting evidence | 87.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OSWorld 2.0 · 2.0 | Agentic | 54.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 · 2.0 | Agentic | 55.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld 2.0 binary without exec · 2.0 | Agentic | 20.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 partial without exec · 2.0 | Agentic | 54.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld Verified · source release snapshot; version not specified | Agentic | 83.4%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld-Verified · source release snapshot; version not specified | Agentic | 83.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| PaperBench · source release snapshot; version not specified | Supporting evidence | 80.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PerceptionBench · source release snapshot; version not specified | Supporting evidence | 47.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PLawBench · source release snapshot; version not specified | Supporting evidence | 69.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 34.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites official PostTrainBench results; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 32.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PRBench-Finance · source release snapshot; version not specified | Supporting evidence | 51.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PRBench-Legal · source release snapshot; version not specified | Supporting evidence | 52.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench · source release snapshot; version not specified | Supporting evidence | 71.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench (Almost Solved) · source release snapshot; version not specified | Supporting evidence | 15.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenQoderBench · source release snapshot; version not specified | Supporting evidence | 62.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenReactBench · source release snapshot; version not specified | Supporting evidence | 1694Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSVGBench · source release snapshot; version not specified | Supporting evidence | 1648Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSWEBench · source release snapshot; version not specified | Supporting evidence | 84%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ResearchRubrics · source release snapshot; version not specified | Supporting evidence | 73.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SaaS-Bench · source release snapshot; version not specified | Supporting evidence | 56.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SciCode · source release snapshot; version not specified | Coding | 53.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SkillsBench · source release snapshot; version not specified | Supporting evidence | 65.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SpreadsheetBench 2 · source release snapshot; version not specified | Supporting evidence | 31.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Bench Pro · not specified | Coding | 69.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 69.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 69.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Marathon · source release snapshot; version not specified | Supporting evidence | 40%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Marathon (v1.1) · v1.1 | Supporting evidence | 48.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 84.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 85%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 85%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 3.0 · 3.0 | Coding | 21.1%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 · 2.1 | Coding | 78.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 84.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 · 2.1 | Coding | 82.7%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon · not specified | Agentic | 59.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / Claude Opus 4.8 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon Verified (Pass@1) · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon-Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon-Verified · source release snapshot; version not specified | Agentic | 76.2%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Video-MME (w. sub) · source release snapshot; version not specified | Supporting evidence | 86%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WebArena Verified · source release snapshot; version not specified | Agentic | 71.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| WideSearch · source release snapshot; version not specified | Supporting evidence | 72.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Item-F1 over4runs; Qwen-Agent for Qwen,ClaudeCode for comparators. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: WideSearch · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WorkSpaceBench · source release snapshot; version not specified | Supporting evidence | 66.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WorldVQA ForceAnswer · source release snapshot; version not specified | Supporting evidence | 39.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ZeroBench (pass@5) · source release snapshot; version not specified | Supporting evidence | 17%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ZeroBench (pass@5) · source release snapshot; version not specified | Supporting evidence | 34%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| τ³-Banking · source release snapshot; version not specified | Supporting evidence | 27.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |