RankingQwen3.5-397B-A17B
Qwen3.5-397B-A17B
42 published benchmark measures · 16 benchmark families contribute across 6 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 3 effort levels across 28 benchmark/harness combinations →
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 46 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
16 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Qwen
- Catalog status
- open-weights-available
- Availability
- Open weights available for self-hosting
- Family
- Qwen3.5
- Released
- —
- Context
- —
- License
- apache-2.0
- Model card
- https://huggingface.co/Qwen/Qwen3.5-397B-A17B
- Default Capability family coverage
/badge/qwen3.5-397b-a17b.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-LCR · source release snapshot; version not specified | Long context | 68.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 68.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: AA-LCR · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AIME 2025 avg@16 · source release snapshot; version not specified | Hard reasoning | 83.1%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AllenAI IFBench · source release snapshot; version not specified | Supporting evidence | 76.5%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Apex-Shortlist (no tools) · source release snapshot; version not specified | Supporting evidence | 61.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 61.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (no tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Apex-Shortlist (with tools) · source release snapshot; version not specified | Supporting evidence | 60.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 60.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Apex-Shortlist (with tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BeyondAIME avg@16 · source release snapshot; version not specified | Supporting evidence | 72.3%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 78.6%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 40.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 40.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: BrowseComp · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Collie · source release snapshot; version not specified | Supporting evidence | 88.9%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CritPt (no tools) · source release snapshot; version not specified | Supporting evidence | 2.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 2.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: CritPt (no tools) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPVal · source release snapshot; version not specified | Agentic | 34.6%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 34.6 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GDPVal · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA (no tools) · source release snapshot; version not specified | Hard reasoning | 87.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 87.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: GPQA (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE (no tools) · source release snapshot; version not specified | Hard reasoning | 28.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 28.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HLE (with tools) · source release snapshot; version not specified | Hard reasoning | 48.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 48.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: HLE (with tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IFBench (prompt loose) · source release snapshot; version not specified | Supporting evidence | 78.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 78.2 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IFBench (prompt loose) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| IMOAnswerBench (no tools) · source release snapshot; version not specified | Hard reasoning | 83.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 83.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (no tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IMOAnswerBench (with tools) · source release snapshot; version not specified | Hard reasoning | 84.51%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 84.51 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IMOAnswerBench (with tools) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IOI 2025 · source release snapshot; version not specified | Supporting evidence | 441.3 pointsReported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 441.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: IOI 2025 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LiveCodeBench (v6) · source release snapshot; version not specified | Coding | 79.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 79.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: LiveCodeBench (v6) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1442 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| Longbench v2 (≤ 1M) · source release snapshot; version not specified | Long context | 68.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 68.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Longbench v2 (≤ 1M) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMLU-Pro | Knowledge | 87.8% | MMLU-Pro reportedcontributes to capability | official board | 2026-09-11 |
| MMLU-Pro · source release snapshot; version not specified | Knowledge | 88.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 88.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specified | Knowledge | 86.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 86.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Multi-Challenge · source release snapshot; version not specified | Supporting evidence | 63.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 63.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Multi-Challenge · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OmniScience Accuracy · source release snapshot; version not specified | Supporting evidence | 35.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 35.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: OmniScience Accuracy · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| PinchBench · source release snapshot; version not specified | Supporting evidence | 86.6%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 86.6 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: PinchBench · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ProfBench (Search) · source release snapshot; version not specified | Supporting evidence | 53%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 53.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: ProfBench (Search) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| RULER (1M) · source release snapshot; version not specified | Long context | 90.1%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 90.1 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: RULER (1M) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SciCode (subtask) · source release snapshot; version not specified | Coding | 48%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 48.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SciCode (subtask) · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Multilingual · source release snapshot; version not specified | Coding | 70.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 70.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Multilingual · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 76.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Verified · source release snapshot; version not specified | Coding | 73.6%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 73.6 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: SWE-Bench Verified · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| tau3 Airline · source release snapshot; version not specified | Supporting evidence | 81.5%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Banking · source release snapshot; version not specified | Supporting evidence | 9.8%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Retail · source release snapshot; version not specified | Supporting evidence | 84.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Telecom · source release snapshot; version not specified | Supporting evidence | 97.8%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Airline · source release snapshot; version not specified | Supporting evidence | 76.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 76.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Airline · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Average · source release snapshot; version not specified | Supporting evidence | 71%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 71.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Average · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Banking · source release snapshot; version not specified | Supporting evidence | 20.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 20.9 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Banking · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Retail · source release snapshot; version not specified | Supporting evidence | 88.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 88.5 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Retail · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TauBench V3 Telecom · source release snapshot; version not specified | Supporting evidence | 98%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 98.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Telecom · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal Bench 2.1 · 2.1 | Coding | 49.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 49.9 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Vals.ai Financial Agent 1.1 with web search · 1.1 | Supporting evidence | 59%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 59.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: with web search · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Vals.ai Financial Agent 1.1 without web search · 1.1 | Supporting evidence | 61.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 61.3 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: without web search · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| WMT24++ (en→xx) · source release snapshot; version not specified | Supporting evidence | 86.8 score (source scale)Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 86.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: WMT24++ (en→xx) · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |