RankingGLM-5.1

GLM-5.1

Data updated 12 Sept 2026

47 published benchmark measures · 16 benchmark families contribute across 6 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GLM-5.1 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GLM-5.1 · Agentic: 16.4 · SupportedGLM-5.1 · Hard reasoning: 9.3 · SupportedGLM-5.1 · Coding: 4.8 · SupportedGLM-5.1 · Human pref: 34.9 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic3 families · 0 with independent evidence · Supported16.4

3 core families; 11 direct opponents across 8 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 12.3–19.6; 0/11 scenarios unsupported. Without one publisher: 14.3–96.3; 0/15 unsupported. Smoothing check: 13.0–24.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 0 with independent evidence · Supported9.3

4 core families; 8 direct opponents across 7 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 7.0–13.2; 0/7 scenarios unsupported. Without one publisher: 2.3–29.8; 1/16 unsupported. Smoothing check: 3.3–21.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding6 families · 0 with independent evidence · Supported4.8

6 core families; 14 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 1.5–23.2; 0/7 scenarios unsupported. Without one publisher: 3.2–20.6; 0/20 unsupported. Smoothing check: 2.7–11.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • SWE-bench Pro: 54.2 observed win share
    Microsoft, MiniMax · Source 1 Source 2
  • Terminal-Bench: 37.5 observed win share
    MiniMax, NVIDIA · Source 1 Source 2
  • SWE-bench Verified: 100.0 observed win share
    Mistral, NVIDIA · Source 1 Source 2
  • LiveCodeBench: 50.0 observed win share
    NVIDIA · Source 1
  • SWE-bench Multilingual: 75.0 observed win share
    NVIDIA · Source 1
  • Scicode: 50.0 observed win share
    NVIDIA · Source 1
Human pref1 families · 1 with independent evidence · Supported34.9

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 26.6–42.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 84.2 observed win share
    lmarena.ai · Source 1
Knowledge1 families · 0 with independent evidence · PreliminaryUnknown

1 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

MultimodalNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context1 families · 0 with independent evidence · PreliminaryUnknown

    1 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Aa Lcr: 25.0 observed win share
      NVIDIA · Source 1

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 4 effort levels across 31 benchmark/harness combinations →

    Reported effort · Not specified

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    16 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Z.ai
    Catalog status
    active
    Availability
    Documented provider API; downloadable official weights
    Family
    GLM
    Released
    Context
    200,000 tokens
    License
    MIT
    Model card
    https://docs.z.ai/guides/overview/overview
    Default Capability family coverage
    Documented-evidence family coverage/badge/glm-5.1.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    AA-LCR · source release snapshot; version not specifiedLong context66.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 66.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AIME 2026 · source release snapshot; version not specifiedHard reasoning95.3%
    Reported settings & source

    MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

    First-party reported result; comparator results retain the source evaluation setup.

    MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Apex-Shortlist (no tools) · source release snapshot; version not specifiedSupporting evidence71.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 71.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Apex-Shortlist (with tools) · source release snapshot; version not specifiedSupporting evidence79%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 79.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic79.3%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic79.3%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic59.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 59.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Claw-Eval · source release snapshot; version not specifiedSupporting evidence62.7%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CritPt (no tools) · source release snapshot; version not specifiedSupporting evidence3.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 3.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPVal · source release snapshot; version not specifiedAgentic54.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 54.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval rubrics · source release snapshot; version not specifiedAgentic68.3%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA (no tools) · source release snapshot; version not specifiedHard reasoning86.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA Diamond · source release snapshot; version not specifiedHard reasoning86.2%
    Reported settings & source

    MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

    First-party reported result; comparator results retain the source evaluation setup.

    MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE (no tools) · source release snapshot; version not specifiedHard reasoning27.2%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 27.2

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE (with tools) · source release snapshot; version not specifiedHard reasoning50.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 50.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HMMT February 2026 · source release snapshot; version not specifiedSupporting evidence82.6%
    Reported settings & source

    MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

    First-party reported result; comparator results retain the source evaluation setup.

    MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    IFBench (prompt loose) · source release snapshot; version not specifiedSupporting evidence76.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 76.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    IMOAnswerBench (no tools) · source release snapshot; version not specifiedHard reasoning86.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IMOAnswerBench (with tools) · source release snapshot; version not specifiedHard reasoning91.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 91.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IOI 2025 · source release snapshot; version not specifiedSupporting evidence456.5 points
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 456.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LiveCodeBench (v6) · source release snapshot; version not specifiedCoding85.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 85.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    LMArena Text ArenaHuman pref1465LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    MCPAtlas · source release snapshot; version not specifiedAgentic71.8%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-Pro · source release snapshot; version not specifiedKnowledge85.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 85.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specifiedKnowledge85.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 85.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Multi-Challenge · source release snapshot; version not specifiedSupporting evidence63%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 63.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    NL2Repo · source release snapshot; version not specifiedSupporting evidence41%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OmniScience Accuracy · source release snapshot; version not specifiedSupporting evidence31.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 31.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PinchBench · source release snapshot; version not specifiedSupporting evidence81.2%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 81.2

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ProfBench (Search) · source release snapshot; version not specifiedSupporting evidence46%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 46.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SciCode (subtask) · source release snapshot; version not specifiedCoding47.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 47.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SpreadsheetBench v1 · source release snapshot; version not specifiedSupporting evidence85.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SVG-Bench · source release snapshot; version not specifiedSupporting evidence56.9%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-Bench Multilingual · source release snapshot; version not specifiedCoding74.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 74.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding58.4%
    Reported settings & source

    MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

    First-party reported result; comparator results retain the source evaluation setup.

    MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Pro · source release snapshot; version not specifiedCoding58.4%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Verified · source release snapshot; version not specifiedCoding80.2%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-Bench Verified · source release snapshot; version not specifiedCoding76.2%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 76.2

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    tau3 Airline · source release snapshot; version not specifiedSupporting evidence79.5%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Banking · source release snapshot; version not specifiedSupporting evidence16.2%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Retail · source release snapshot; version not specifiedSupporting evidence76.3%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Telecom · source release snapshot; version not specifiedSupporting evidence98.7%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Airline · source release snapshot; version not specifiedSupporting evidence85%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 85.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Average · source release snapshot; version not specifiedSupporting evidence69.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 69.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Banking · source release snapshot; version not specifiedSupporting evidence12.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 12.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Retail · source release snapshot; version not specifiedSupporting evidence84.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 84.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Telecom · source release snapshot; version not specifiedSupporting evidence96.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 96.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Terminal Bench 2.1 · 2.1Coding59.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 59.3 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.0 · 2.0Coding69%
    Reported settings & source

    MAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports.

    First-party reported result; comparator results retain the source evaluation setup.

    MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Microsoft report Table 11 and section 4.1: MAI Terminal-Bench removes timeouts and uses a minimal ReAct harness; comparator values are cited from other model releases. A common evaluation protocol is not established.

    Reviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding48.3%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Vals.ai Financial Agent 1.1 with web search · 1.1Supporting evidence60.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 60.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Vals.ai Financial Agent 1.1 without web search · 1.1Supporting evidence60.2%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 60.2

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    VIBE-V2 · source release snapshot; version not specifiedSupporting evidence48.2%
    Reported settings & source

    MiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted.

    First-party reported result; comparator results retain the source evaluation setup.

    MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    WMT24++ (en→xx) · source release snapshot; version not specifiedSupporting evidence84.4 score (source scale)
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 84.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ARC-AGI-2Hard reasoning
    DeepSWE v1.1Agentic
    GDPval-AAAgentic
    GPQA DiamondHard reasoning
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    OSWorld-VerifiedAgentic
    SWE-bench ProAgentic
    SWE-bench VerifiedAgentic
    Terminal-Bench 2.1Agentic