RankingLVBench

LVBench

Data updated 30 Sept 2026

Bucket
Multimodal
Unit
percent
Direction
Higher is better
Version
—
Display harness
Google Gemini 4 Argon launch and evaluation methodology
Board
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–4 of 4 entries

ModelScoreHarnessEvidenceSource-recorded date
Gemini 4 ArgonGoogle91.7%
Reported settings & source

Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison.

Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

Google Gemini 4 Argon launch and evaluation methodologylab self-report2026-09-30
Claude Fable 5.1Anthropic79.7%
Reported settings & source

Comparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison.

Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / Claude Fable 5.1; methodology Additional Details · reviewed 2026-09-30

Google Gemini 4 Argon launch and evaluation methodologylab self-report2026-09-30
Claude Opus 5.5Anthropic83.7%
Reported settings & source

Comparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison.

Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / Claude Opus 5.5; methodology Additional Details · reviewed 2026-09-30

Google Gemini 4 Argon launch and evaluation methodologylab self-report2026-09-30
GPT-6 AstraOpenAI87.5%
Reported settings & source

Comparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison.

Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / GPT-6 Astra; methodology Additional Details · reviewed 2026-09-30

Google Gemini 4 Argon launch and evaluation methodologylab self-report2026-09-30