RankingMMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specified
MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · source release snapshot; version not specified
- Bucket
- Knowledge
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- Board
- https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Qwen3.5-397B-A17BQwen | 86.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 86.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 85.8%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 85.8 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| Kimi-K2.6Moonshot | 85%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 85.0 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| NVIDIA Nemotron 3 UltraNVIDIA | 83%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 83.0 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| MiniMax-M2.7MiniMax | 78.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 78.4 nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-ProX (avg en/de/fr/es/it/ja/zh/hi/pt/ko) · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |