RankingHarvey Legal Agent Benchmark

Harvey Legal Agent Benchmark

Data updated 21 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
Display harness
Grok 4.7 launch comparison
Board
https://x.ai/news/grok-4-7

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–4 of 4 entries

ModelScoreHarnessEvidenceSource-recorded date
Grok 4.7xAI19.6%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab

Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
Grok 4.6xAI15.8%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab

Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
Claude Fable 5.1Anthropic6.7%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab

Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21
GPT-5.6 SolOpenAI2.5%
Reported settings & source

Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab

Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21

Grok 4.7 launch comparisonlab self-report2026-09-21