RankingOCRBench v2 average accuracy · source release snapshot; version not specified

OCRBench v2 average accuracy · source release snapshot; version not specified

Data updated 12 Sept 2026

Bucket
Multimodal
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
Amazon Nova 2 technical report
Board
https://cdn.amazon.science/c5/3d/84514a224666b5be6de4b43ef4aa/nova-2-0-technical-report2.pdf

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Amazon Nova 2 Pro PreviewAmazon64.5%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova PremierAmazon61.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova ProAmazon59.6%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Gemini 2.5 ProGoogle59.3%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Gemini 2.5 FlashGoogle58.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5OpenAI58.1%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5.1OpenAI57.2%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova 2 LiteAmazon56.1%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
GPT-5 MiniOpenAI55.4%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Amazon Nova LiteAmazon55.1%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Sonnet 4.5Anthropic52.3%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06
Claude Haiku 4.5Anthropic48.3%
Reported settings & source

Nova2 launch evaluation; benchmark-specific settings in report. tau2 Verified avg@3; telecom cited AA because user model sensitivity; comparators mix own runs and cited provider scores.

First-party reported result; comparator results retain the source evaluation setup.

Amazon Nova 2 technical report · Table3 · reviewed 2026-09-06

Amazon Nova 2 technical reportlab self-report2026-09-06