RankingGLM-5.3

GLM-5.3

Data updated 12 Sept 2026

18 published benchmark measures · 6 benchmark families contribute across 3 task areas. 3 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GLM-5.3 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GLM-5.3 · Agentic: 68.6 · SupportedGLM-5.3 · Hard reasoning: 64.2 · SupportedGLM-5.3 · Coding: 69.6 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic2 families · 0 with independent evidence · Supported68.6

2 core families; 6 direct opponents across 6 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 65.2–69.9; 2/11 scenarios unsupported. Without one publisher: 65.7–71.3; 1/15 unsupported. Smoothing check: 64.1–71.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • AutomationBench: 100.0 observed win share
    Z.ai · Source 1
  • Toolathlon: 33.3 observed win share
    Z.ai · Source 1
Hard reasoning2 families · 1 with independent evidence · Supported64.2

2 core families; 25 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 60.0–65.3; 2/7 scenarios unsupported. Without one publisher: 58.0–66.7; 2/16 unsupported. Smoothing check: 60.9–65.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GPQA: 31.8 observed win share
    artificialanalysis.ai · Source 1
  • Humanity’s Last Exam: 83.3 observed win share
    Z.ai · Source 1
Coding2 families · 1 with independent evidence · Supported69.6

2 core families; 27 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 65.5–70.3; 2/7 scenarios unsupported. Without one publisher: 66.9–71.3; 1/20 unsupported. Smoothing check: 65.4–71.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      MultimodalNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Long contextNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

          Compare 3 effort levels across 36 benchmark/harness combinations →

          Reported effort · Max + unspecified

          Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

          Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

          Inspect each result and its source ↓ · Download effort evidence

          Compare capability profiles →

          Score contributions and missing evidence

          6 contributing families across 3 capabilities. Fixed reference panels do not change when the catalog expands.

          Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

          Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

          Model information & shareable badge
          Lab
          Z.ai
          Catalog status
          active
          Availability
          Documented provider API; downloadable official weights
          Family
          GLM
          Released
          Context
          1,000,000 tokens
          License
          Model card
          https://docs.z.ai/guides/overview/overview
          Default Capability family coverage
          Documented-evidence family coverage/badge/glm-5.3.svg
          Benchmark scores & sources

          Original results, evaluation harnesses, and evidence behind this model.

          BenchmarkBucketScoreHarnessEvidenceSource-recorded date
          Agents' Last Exam (ALE-CLI) · source release snapshot; version not specifiedSupporting evidence28.5%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          AutomationBench (v1.0.6) · v1.0.6Agentic48.2%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          contributes to capability
          lab self-reportReviewed 2026-09-12
          CyberGym · source release snapshot; version not specifiedSupporting evidence84.5%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          DeepSWE (v1.1) · v1.1Coding66.9%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          contributes to capability
          lab self-reportReviewed 2026-09-12
          DeepSWE v1.1Coding69%DeepSWE v1.1 reported

          Effort: Not specified

          contributes to capability
          official board2026-09-03
          ExploitBench · source release snapshot; version not specifiedSupporting evidence54.4%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence105 count
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          ExploitGym (2h / 6h) · source release snapshot; version not specifiedSupporting evidence130 count
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          FrontierSWE · source release snapshot; version not specifiedSupporting evidence78.1%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          Externally evaluated result, attributed in the source footnote.

          Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison.

          zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          GDPval-AA v2 · source release snapshot; version not specifiedAgentic1769
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          Externally evaluated result, attributed in the source footnote.

          Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

          zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          lab self-report

          The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison.

          Reviewed 2026-09-12
          GPQA DiamondHard reasoning91.717%GPQA Diamond reported

          Effort: Not specified

          contributes to capability
          independent repro2026-09-11
          HLE w/ Tools · source release snapshot; version not specifiedHard reasoning62.5%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          contributes to capability
          lab self-reportReviewed 2026-09-12
          NL2Repo · source release snapshot; version not specifiedSupporting evidence58%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          PostTrainBench · source release snapshot; version not specifiedSupporting evidence39.8%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          Comparison limit: The source changes anti-cheat checks and substitutes zero-shot baseline scores for failed runs; the modified protocol requires separate review.

          zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          ProgramBench (Almost Solved) · source release snapshot; version not specifiedSupporting evidence19%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          SWE-Marathon (v1.1) · v1.1Supporting evidence42.5%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          Comparison limit: The source replaces affected anti-cheat checks with LLM inspection and changes task images; the modified protocol requires separate review.

          zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12

          Published configuration

          Effort: Max

          lab self-report

          Supporting evidence outside the reviewed capability core

          Reviewed 2026-09-12
          Terminal Bench 2.1 · 2.1Coding88.2%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          contributes to capability
          lab self-reportReviewed 2026-09-12
          Terminal Bench 3.0 · 3.0Coding28.3%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12

          Published configuration

          Effort: Max

          contributes to capability
          lab self-reportReviewed 2026-09-12
          Toolathlon Verified · source release snapshot; version not specifiedAgentic73%
          Reported settings & source

          GLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details.

          First-party reported result.

          zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12

          Published configuration

          Effort: Not specified

          contributes to capability
          lab self-reportReviewed 2026-09-12
          VulcanBench v3Supporting evidence73.9%VulcanBench v3 bare-bones API · Report 18 · high

          Effort: Not specified

          official board

          Supporting evidence outside the reviewed capability core

          2026-08-24
          VulcanBench v3Supporting evidence78.3%VulcanBench v3 bare-bones API · Report 18 · low

          Effort: Not specified

          official board

          Supporting evidence outside the reviewed capability core

          2026-08-24
          VulcanBench v3Supporting evidence65.2%VulcanBench v3 bare-bones API · Report 18 · max

          Effort: Not specified

          official board

          Supporting evidence outside the reviewed capability core

          2026-08-24
          ARC-AGI-2Hard reasoning
          GDPval-AAAgentic
          Humanity's Last ExamHard reasoning
          LiveCodeBenchCoding
          LMArena Text ArenaHuman pref
          MMLU-ProKnowledge
          OSWorld-VerifiedAgentic
          SWE-bench ProAgentic
          SWE-bench VerifiedAgentic
          Terminal-Bench 2.1Agentic