RankingGrok 4.7
Grok 4.7
8 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 1 observations
- XHigh: 7 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- xAI
- Catalog status
- active
- Availability
- Available through the Grok API, Grok Build, Cursor, and supported third-party platforms; account and region restrictions may apply
- Family
- Grok 4
- Released
- 2026-09-21
- Context
- 500,000 tokens
- License
- proprietary
- Model card
- https://docs.x.ai/developers/models/grok-4.7
- Default Capability family coverage
/badge/grok-4.7.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA Briefcase · 1.1 | Supporting evidence | 1657Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons. Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| CursorBench · 4.0 | Supporting evidence | 46.3%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: CursorBench · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| DeepSWE v1.1 · 1.1 | Coding | 71%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: DeepSWE v1.1 · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| EEBench | Supporting evidence | 64%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: EEBench · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| GDPval-AA · 2 | Agentic | 1695Reported settings & sourceGrok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart. Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence. Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| Harvey Legal Agent Benchmark | Supporting evidence | 19.6%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings. Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| HealthBench Professional | Supporting evidence | 56.7%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: HealthBench Professional · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| Terminal-Bench · 4.0 | Coding | 38%Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation. Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21 | Published configuration | lab self-report | Reviewed 2026-09-21 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |