RankingQwen3.5-397B-A17B

Qwen3.5-397B-A17B

Data updated 12 Sept 2026

42 published benchmark measures · 16 benchmark families contribute across 6 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Qwen3.5-397B-A17B capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Qwen3.5-397B-A17B · Agentic: 1.6 · SupportedQwen3.5-397B-A17B · Hard reasoning: 4.5 · SupportedQwen3.5-397B-A17B · Coding: 3.2 · SupportedQwen3.5-397B-A17B · Human pref: 22.7 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic2 families · 0 with independent evidence · Supported1.6

2 core families; 7 direct opponents across 6 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 1.3–2.1; 2/11 scenarios unsupported. Without one publisher: 0.6–52.6; 1/15 unsupported. Smoothing check: 0.5–6.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning4 families · 0 with independent evidence · Supported4.5

4 core families; 8 direct opponents across 6 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 3.3–7.6; 0/7 scenarios unsupported. Without one publisher: 1.6–21.4; 1/16 unsupported. Smoothing check: 1.2–14.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding5 families · 0 with independent evidence · Supported3.2

5 core families; 7 direct opponents across 6 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 1.1–21.2; 0/7 scenarios unsupported. Without one publisher: 2.1–15.2; 1/20 unsupported. Smoothing check: 1.6–8.5; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • SWE-bench Verified: 12.5 observed win share
    Mistral, NVIDIA · Source 1 Source 2
  • LiveCodeBench: 25.0 observed win share
    NVIDIA · Source 1
  • SWE-bench Multilingual: 25.0 observed win share
    NVIDIA · Source 1
  • Scicode: 75.0 observed win share
    NVIDIA · Source 1
  • Terminal-Bench: 0.0 observed win share
    NVIDIA · Source 1
Human pref1 families · 1 with independent evidence · Supported22.7

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 12.4–34.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 73.7 observed win share
    lmarena.ai · Source 1
Knowledge1 families · 1 with independent evidence · PreliminaryUnknown

1 core families; 55 direct opponents across 14 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

MultimodalNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context3 families · 0 with independent evidence · PreliminaryUnknown

    3 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Aa Lcr: 50.0 observed win share
      NVIDIA · Source 1
    • Longbench: 100.0 observed win share
      NVIDIA · Source 1
    • Ruler 1M: 0.0 observed win share
      NVIDIA · Source 1

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 3 effort levels across 28 benchmark/harness combinations →

    Reported effort · Not specified

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    16 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Qwen
    Catalog status
    open-weights-available
    Availability
    Open weights available for self-hosting
    Family
    Qwen3.5
    Released
    Context
    License
    apache-2.0
    Model card
    https://huggingface.co/Qwen/Qwen3.5-397B-A17B
    Default Capability family coverage
    Documented-evidence family coverage/badge/qwen3.5-397b-a17b.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    AA-LCR · source release snapshot; version not specifiedLong context68.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 68.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AIME 2025 avg@16 · source release snapshot; version not specifiedHard reasoning83.1%
    Reported settings & source

    Maximum reasoning; Sonnet4.6 external API truncation caveat.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AllenAI IFBench · source release snapshot; version not specifiedSupporting evidence76.5%
    Reported settings & source

    Maximum reasoning; Sonnet4.6 external API truncation caveat.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Apex-Shortlist (no tools) · source release snapshot; version not specifiedSupporting evidence61.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 61.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Apex-Shortlist (with tools) · source release snapshot; version not specifiedSupporting evidence60.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 60.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BeyondAIME avg@16 · source release snapshot; version not specifiedSupporting evidence72.3%
    Reported settings & source

    Maximum reasoning; Sonnet4.6 external API truncation caveat.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic78.6%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BrowseComp · source release snapshot; version not specifiedAgentic40.5%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 40.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Collie · source release snapshot; version not specifiedSupporting evidence88.9%
    Reported settings & source

    Maximum reasoning; Sonnet4.6 external API truncation caveat.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CritPt (no tools) · source release snapshot; version not specifiedSupporting evidence2.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 2.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPVal · source release snapshot; version not specifiedAgentic34.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 34.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA (no tools) · source release snapshot; version not specifiedHard reasoning87.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 87.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE (no tools) · source release snapshot; version not specifiedHard reasoning28.5%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 28.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE (with tools) · source release snapshot; version not specifiedHard reasoning48.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 48.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IFBench (prompt loose) · source release snapshot; version not specifiedSupporting evidence78.2%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 78.2

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    IMOAnswerBench (no tools) · source release snapshot; version not specifiedHard reasoning83.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 83.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IMOAnswerBench (with tools) · source release snapshot; version not specifiedHard reasoning84.51%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 84.51

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IOI 2025 · source release snapshot; version not specifiedSupporting evidence441.3 points
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 441.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LiveCodeBench (v6) · source release snapshot; version not specifiedCoding79.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 79.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    LMArena Text ArenaHuman pref1442LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    Longbench v2 (≤ 1M) · source release snapshot; version not specifiedLong context68.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 68.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Longbench v2 (≤ 1M) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-ProKnowledge87.8%MMLU-Pro reported

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    MMLU-Pro · source release snapshot; version not specifiedKnowledge88.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 88.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specifiedKnowledge86.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Multi-Challenge · source release snapshot; version not specifiedSupporting evidence63.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 63.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OmniScience Accuracy · source release snapshot; version not specifiedSupporting evidence35.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 35.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PinchBench · source release snapshot; version not specifiedSupporting evidence86.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ProfBench (Search) · source release snapshot; version not specifiedSupporting evidence53%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 53.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    RULER (1M) · source release snapshot; version not specifiedLong context90.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 90.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: RULER (1M) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SciCode (subtask) · source release snapshot; version not specifiedCoding48%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 48.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-Bench Multilingual · source release snapshot; version not specifiedCoding70.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 70.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench Verified · source release snapshot; version not specifiedCoding76.4%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-Bench Verified · source release snapshot; version not specifiedCoding73.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 73.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    tau3 Airline · source release snapshot; version not specifiedSupporting evidence81.5%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Banking · source release snapshot; version not specifiedSupporting evidence9.8%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Retail · source release snapshot; version not specifiedSupporting evidence84.4%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    tau3 Telecom · source release snapshot; version not specifiedSupporting evidence97.8%
    Reported settings & source

    Maximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k.

    First-party reported result; comparator results retain the source evaluation setup.

    Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Airline · source release snapshot; version not specifiedSupporting evidence76.5%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 76.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Average · source release snapshot; version not specifiedSupporting evidence71%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 71.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Banking · source release snapshot; version not specifiedSupporting evidence20.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 20.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Retail · source release snapshot; version not specifiedSupporting evidence88.5%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 88.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Telecom · source release snapshot; version not specifiedSupporting evidence98%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 98.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Terminal Bench 2.1 · 2.1Coding49.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 49.9 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Vals.ai Financial Agent 1.1 with web search · 1.1Supporting evidence59%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 59.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Vals.ai Financial Agent 1.1 without web search · 1.1Supporting evidence61.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 61.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    WMT24++ (en→xx) · source release snapshot; version not specifiedSupporting evidence86.8 score (source scale)
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ARC-AGI-2Hard reasoning
    DeepSWE v1.1Agentic
    GDPval-AAAgentic
    GPQA DiamondHard reasoning
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    OSWorld-VerifiedAgentic
    SWE-bench ProAgentic
    SWE-bench VerifiedAgentic
    Terminal-Bench 2.1Agentic