RankingGLM-5.2
GLM-5.2
45 published benchmark measures · 9 benchmark families contribute across 3 task areas. 3 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 7 effort levels across 44 benchmark/harness combinations →
Reported effort · Max + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Max: 28 observations
- Not specified: 29 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
9 contributing families across 3 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Z.ai
- Catalog status
- active
- Availability
- Documented provider API; downloadable official weights
- Family
- GLM
- Released
- —
- Context
- 1,000,000 tokens
- License
- MIT
- Model card
- https://docs.z.ai/guides/overview/overview
- Default Capability family coverage
/badge/glm-5.2.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase (Elo) · source release snapshot; version not specified | Supporting evidence | 1260Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-LCR · source release snapshot; version not specified | Long context | 71.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam · source release snapshot; version not specified | Supporting evidence | 20.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam · source release snapshot; version not specified | Supporting evidence | 23.8%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (ALE-CLI) · source release snapshot; version not specified | Supporting evidence | 23.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 35.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-2 | Hard reasoning | 22.8% | ARC Prize verifiedcontributes to capability | official board | 2026-06-13 |
| AutomationBench · source release snapshot; version not specified | Agentic | 12.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (Public) · source release snapshot; version not specified | Agentic | 12.9%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: AutomationBench (Public) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (v1.0.6) · v1.0.6 | Agentic | 26.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CorpFin v2 · source release snapshot; version not specified | Supporting evidence | 66.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CritPt · source release snapshot; version not specified | Supporting evidence | 20.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CyberGym · source release snapshot; version not specified | Supporting evidence | 77.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSWE · source release snapshot; version not specified | Coding | 46.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE · source release snapshot; version not specified | Coding | 46.2%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DeepSWE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE (v1.1) · v1.1 | Coding | 46.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 43.8% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DSBench-FullStack † · source release snapshot; version not specified | Supporting evidence | 61.8%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-FullStack † · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DSBench-Hard † · source release snapshot; version not specified | Supporting evidence | 54.5%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. † source footnote applies. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: DSBench-Hard † · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitBench · source release snapshot; version not specified | Supporting evidence | 24.4%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 29 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 39 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 49.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 67.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 67.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 · source release snapshot; version not specified | Agentic | 1508Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 (Elo) · source release snapshot; version not specified | Agentic | 1510Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 91.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Harvey Lab-AA · source release snapshot; version not specified | Supporting evidence | 91%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HLE (wo / w tools) · source release snapshot; version not specified | Hard reasoning | 40.5%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. without tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE (wo / w tools) · source release snapshot; version not specified | Hard reasoning | 54.7%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. with tools First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: HLE (wo / w tools) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ Tools · source release snapshot; version not specified | Hard reasoning | 54.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 43.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Kimi Code Bench 2.0 · 2.0 | Supporting evidence | 64.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Legal Research Bench · source release snapshot; version not specified | Supporting evidence | 31.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MCP-Atlas · source release snapshot; version not specified | Agentic | 82.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 40.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 48.9%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: NL2Repo · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 48.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OfficeQA Pro · source release snapshot; version not specified | Agentic | 41.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 34.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites official PostTrainBench results; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 31.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench · source release snapshot; version not specified | Supporting evidence | 63.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench (Almost Solved) · source release snapshot; version not specified | Supporting evidence | 9.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ResearchRubrics · source release snapshot; version not specified | Supporting evidence | 71.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SciCode · source release snapshot; version not specified | Coding | 50.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SpreadsheetBench 2 · source release snapshot; version not specified | Supporting evidence | 28.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Marathon · source release snapshot; version not specified | Supporting evidence | 13%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM-5.2 release blog; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Marathon (v1.1) · v1.1 | Supporting evidence | 19.4%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 81%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 81%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 3.0 · 3.0 | Coding | 4.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 · 2.1 | Coding | 82.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 59.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon-Verified · source release snapshot; version not specified | Agentic | 59.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon-Verified · source release snapshot; version not specified | Agentic | 59.9%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| τ³-Banking · source release snapshot; version not specified | Supporting evidence | 26.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |