RankingNVIDIA Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra

Data updated 12 Sept 2026

41 published benchmark measures · 15 benchmark families contribute across 6 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
NVIDIA Nemotron 3 Ultra capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100NVIDIA Nemotron 3 Ultra · Agentic: 2.7 · SupportedNVIDIA Nemotron 3 Ultra · Hard reasoning: 15.0 · SupportedNVIDIA Nemotron 3 Ultra · Coding: 1.7 · SupportedNVIDIA Nemotron 3 Ultra · Human pref: 15.7 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic2 families · 1 with independent evidence · Supported2.7

2 core families; 43 direct opponents across 18 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 2.3–3.6; 2/11 scenarios unsupported. Without one publisher: 1.7–72.7; 1/15 unsupported. Smoothing check: 1.1–8.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 47.5 observed win share
    NVIDIA, artificialanalysis.ai · Source 1 Source 2
  • Browsecomp: 25.0 observed win share
    NVIDIA · Source 1
Hard reasoning3 families · 1 with independent evidence · Supported15.0

3 core families; 26 direct opponents across 11 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 5.6–24.1; 0/7 scenarios unsupported. Without one publisher: 4.9–41.8; 1/16 unsupported. Smoothing check: 6.5–28.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding5 families · 0 with independent evidence · Supported1.7

5 core families; 4 direct opponents across 4 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 0.5–14.3; 0/7 scenarios unsupported. Without one publisher: 1.0–9.2; 1/20 unsupported. Smoothing check: 0.7–5.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • LiveCodeBench: 75.0 observed win share
    NVIDIA · Source 1
  • SWE-bench Multilingual: 0.0 observed win share
    NVIDIA · Source 1
  • SWE-bench Verified: 0.0 observed win share
    NVIDIA · Source 1
  • Scicode: 25.0 observed win share
    NVIDIA · Source 1
  • Terminal-Bench: 50.0 observed win share
    NVIDIA · Source 1
Human pref1 families · 1 with independent evidence · Supported15.7

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 6.6–28.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 64.9 observed win share
    lmarena.ai · Source 1
Knowledge1 families · 0 with independent evidence · PreliminaryUnknown

1 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

MultimodalNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Long context3 families · 0 with independent evidence · PreliminaryUnknown

    3 core families; 4 direct opponents across 4 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Aa Lcr: 0.0 observed win share
      NVIDIA · Source 1
    • Longbench: 0.0 observed win share
      NVIDIA · Source 1
    • Ruler 1M: 100.0 observed win share
      NVIDIA · Source 1

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 2 effort levels across 27 benchmark/harness combinations →

    Reported effort · Thinking + unspecified

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    15 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    NVIDIA
    Catalog status
    open-weights-available
    Availability
    Open weights available for self-hosting
    Family
    Nemotron 3
    Released
    Context
    License
    Model card
    https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
    Default Capability family coverage
    Documented-evidence family coverage/badge/nemotron-3-ultra.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    AA-LCR · source release snapshot; version not specifiedLong context65.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 65.4 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Apex-Shortlist (no tools) · source release snapshot; version not specifiedSupporting evidence74.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 74.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Apex-Shortlist (with tools) · source release snapshot; version not specifiedSupporting evidence84.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 84.8

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence IndexSupporting evidence38 AA Intelligence Index

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-09-02
    BrowseCompAgentic44.4%BrowseComp reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    BrowseComp · source release snapshot; version not specifiedAgentic44.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 44.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    CritPt (no tools) · source release snapshot; version not specifiedSupporting evidence3.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 3.1 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPVal · source release snapshot; version not specifiedAgentic46.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 46.7 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AAAgentic1091Artificial Analysis GDPval-AA

    Effort: Not specified

    contributes to capability
    official board2026-09-12
    GPQA (no tools) · source release snapshot; version not specifiedHard reasoning87%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 87.0 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GPQA DiamondHard reasoning86.667%GPQA Diamond reported

    Effort: Not specified

    contributes to capability
    independent repro2026-09-11
    HLE (no tools) · source release snapshot; version not specifiedHard reasoning26.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 26.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE (with tools) · source release snapshot; version not specifiedHard reasoning37.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 37.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Humanity's Last ExamHard reasoning26.7%HLE no tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    IFBench (prompt loose) · source release snapshot; version not specifiedSupporting evidence81.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 81.7 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    IMOAnswerBench (no tools) · source release snapshot; version not specifiedHard reasoning88.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 88.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IMOAnswerBench (with tools) · source release snapshot; version not specifiedHard reasoning92.3%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 92.3

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    IOI 2025 · source release snapshot; version not specifiedSupporting evidence570 points
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 570.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LiveCodeBenchCoding89%LiveCodeBench pass@1

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    LiveCodeBench (v6) · source release snapshot; version not specifiedCoding89%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 89.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    LMArena Text ArenaHuman pref1426LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    Longbench v2 (≤ 1M) · source release snapshot; version not specifiedLong context61.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 61.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Longbench v2 (≤ 1M) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-ProKnowledge86.8%MMLU-Pro reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    MMLU-Pro · source release snapshot; version not specifiedKnowledge86.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.8 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specifiedKnowledge83%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 83.0 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Multi-Challenge · source release snapshot; version not specifiedSupporting evidence63.8%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 63.8 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    OmniScience Accuracy · source release snapshot; version not specifiedSupporting evidence24.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 24.1 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    PinchBench · source release snapshot; version not specifiedSupporting evidence90%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 90.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ProfBench (Search) · source release snapshot; version not specifiedSupporting evidence56%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 56.0

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    RULER (1M) · source release snapshot; version not specifiedLong context94.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 94.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: RULER (1M) · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SciCode (subtask) · source release snapshot; version not specifiedCoding44.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 44.6 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-Bench Multilingual · source release snapshot; version not specifiedCoding67.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 67.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    SWE-bench VerifiedCoding71.9%SWE-bench Verified official

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    SWE-Bench Verified · source release snapshot; version not specifiedCoding70.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 70.7 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    contributes to capability
    lab self-reportReviewed 2026-09-06
    TauBench V3 Airline · source release snapshot; version not specifiedSupporting evidence81.5%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 81.5

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Average · source release snapshot; version not specifiedSupporting evidence70.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 70.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Banking · source release snapshot; version not specifiedSupporting evidence22.6%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 22.6

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Retail · source release snapshot; version not specifiedSupporting evidence86.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 86.4

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    TauBench V3 Telecom · source release snapshot; version not specifiedSupporting evidence92.9%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 92.9

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Terminal Bench 2.1 · 2.1Coding56.4%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 56.4 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1Coding56.4%Terminal-Bench 2.1 reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-06-04
    Vals.ai Financial Agent 1.1 with web search · 1.1Supporting evidence53.7%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 53.7

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Vals.ai Financial Agent 1.1 without web search · 1.1Supporting evidence60.1%
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 60.1

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    WMT24++ (en→xx) · source release snapshot; version not specifiedSupporting evidence83.7 score (source scale)
    Reported settings & source

    NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

    First-party reported result. Original cell: 83.7 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

    nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06

    Published configuration

    Effort: Thinking

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    τ²-bench TelecomSupporting evidence92.9%τ²-bench Telecom reported

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    2026-06-04
    ARC-AGI-2Hard reasoning
    DeepSWE v1.1Agentic
    OSWorld-VerifiedAgentic
    SWE-bench ProAgentic