RankingHarvey's Legal Agent Benchmark · 1
Harvey's Legal Agent Benchmark · 1
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- 1
- Display harness
- Google Gemini 4 Argon launch and evaluation methodology
- Board
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Gemini 4 ArgonGoogle | 19.6%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Google Gemini 4 Argon launch and evaluation methodology | lab self-report | 2026-09-30 |
| Claude Fable 5.1Anthropic | 6.7%Reported settings & sourceComparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / Claude Fable 5.1; methodology Additional Details · reviewed 2026-09-30 | Google Gemini 4 Argon launch and evaluation methodology | lab self-report | 2026-09-30 |
| Claude Opus 5.5Anthropic | 3.8%Reported settings & sourceComparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / Claude Opus 5.5; methodology Additional Details · reviewed 2026-09-30 | Google Gemini 4 Argon launch and evaluation methodology | lab self-report | 2026-09-30 |
| GPT-6 AstraOpenAI | 5.4%Reported settings & sourceComparator maximum available thinking/reasoning when reported, otherwise best available result; exact setting not identified by this grid. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / GPT-6 Astra; methodology Additional Details · reviewed 2026-09-30 | Google Gemini 4 Argon launch and evaluation methodology | lab self-report | 2026-09-30 |