RankingMMLU-Pro · source release snapshot; version not specified

MMLU-Pro · source release snapshot; version not specified

Data updated 12 Sept 2026

Bucket
Knowledge
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
Board
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Qwen3.5-397B-A17BQwen88.3%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 88.3

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16lab self-report2026-09-06
Kimi-K2.6Moonshot88.1%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 88.1

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16lab self-report2026-09-06
NVIDIA Nemotron 3 UltraNVIDIA86.8%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 86.8 The linked NVIDIA reproduction recipe explicitly enables thinking for this model and benchmark. It does not establish Low/Medium/High/Max or a fixed reasoning-token budget; the request adapter removes max_tokens and max_completion_tokens. Comparator settings are not specified by this recipe.

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16lab self-report2026-09-06
GLM-5.1Z.ai85.9%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 85.9

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16lab self-report2026-09-06
MiniMax-M2.7MiniMax81.9%
Reported settings & source

NVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol.

First-party reported result. Original cell: 81.9

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: MMLU-Pro · reviewed 2026-09-06

nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16lab self-report2026-09-06
Amazon Nova 2 LiteAmazon80.9%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova 2 Pro PreviewAmazon81.6%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova LiteAmazon59%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova PremierAmazon73.3%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova ProAmazon69.1%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Haiku 4.5Anthropic80%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Sonnet 4.5Anthropic87.5%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Gemini 2.5 FlashGoogle83.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Gemini 2.5 ProGoogle86.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5OpenAI87.1%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5 MiniOpenAI83.7%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5.1OpenAI87%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06