RankingAmazon Nova 2 Pro Preview
Amazon Nova 2 Pro Preview
18 published benchmark measures · 10 benchmark families contribute across 5 task areas. 3 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 3 effort levels across 17 benchmark/harness combinations →
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 19 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
10 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Amazon
- Catalog status
- availability-unverified
- Availability
- Announced preview for Nova Forge early-access customers; current general API availability not confirmed
- Family
- Nova
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://aws.amazon.com/about-aws/whats-new/2025/12/nova-2-foundation-models-amazon-bedrock/
- Default Capability family coverage
/badge/amazon-nova-2-pro-preview.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AIME 2025 · source release snapshot; version not specified | Hard reasoning | 92.3%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BFCL v4 · source release snapshot; version not specified | Agentic | 61.6%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 81.4%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IFBench prompt loose · source release snapshot; version not specified | Supporting evidence | 80.2%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LiveCodeBench v5 July2024-January2025 · source release snapshot; version not specified | Coding | 74.6%Reported settings & sourceNova2 internal agentic scaffold; scaled inference explicitly separate. Comparator setups from cited provider report. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table4 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| LongCodeBench 1M · source release snapshot; version not specified | Supporting evidence | 84%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MMLU-Pro · source release snapshot; version not specified | Knowledge | 81.6%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 63.5%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MultiChallenge · source release snapshot; version not specified | Supporting evidence | 77.7%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OCRBench v2 average accuracy · source release snapshot; version not specified | Multimodal | 64.5%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| QVHighlights R1@0.5 · 0.5 | Supporting evidence | 76.7%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| RealKIE-FCC Verified ALNS · source release snapshot; version not specified | Supporting evidence | 67%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ScreenSpot point accuracy · source release snapshot; version not specified | Multimodal | 88.1%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 61.5%Reported settings & sourceNova2 internal agentic scaffold; scaled inference explicitly separate. Comparator setups from cited provider report. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table4 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified scaled inference · source release snapshot; version not specified | Coding | 70%Reported settings & sourceNova2 internal agentic scaffold; scaled inference explicitly separate. Comparator setups from cited provider report. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table4 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| tau2 Airline Verified · source release snapshot; version not specified | Supporting evidence | 65.2%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau2 Retail Verified · source release snapshot; version not specified | Supporting evidence | 77.7%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau2 Telecom · source release snapshot; version not specified | Supporting evidence | 92.7%Reported settings & sourceNova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table2 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 1.0 · 1.0 | Coding | 41.3%Reported settings & sourceNova2 internal agentic scaffold; scaled inference explicitly separate. Comparator setups from cited provider report. First-party reported result; comparator results retain the source evaluation setup. Amazon Nova 2 technical report · Table4 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |