RankingHarvey Legal Agent Benchmark
Harvey Legal Agent Benchmark
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- —
- Display harness
- Grok 4.7 launch comparison
- Board
- https://x.ai/news/grok-4-7
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Grok 4.7xAI | 19.6%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings. Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.6xAI | 15.8%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings. Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Claude Fable 5.1Anthropic | 6.7%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings. Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| GPT-5.6 SolOpenAI | 2.5%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings. Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |