RankingGrok 4.7

Grok 4.7

Compare models

Data updated 21 Sept 2026

8 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Grok 4.7 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 20/20 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Mixed settings

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                xAI
                Catalog status
                active
                Availability
                Available through the Grok API, Grok Build, Cursor, and supported third-party platforms; account and region restrictions may apply
                Family
                Grok 4
                Released
                2026-09-21
                Context
                500,000 tokens
                License
                proprietary
                Model card
                https://docs.x.ai/developers/models/grok-4.7
                Default Capability family coverage
                Documented-evidence family coverage/badge/grok-4.7.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                AA Briefcase · 1.1Supporting evidence1657
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase

                Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons.

                Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons.

                Reviewed 2026-09-21
                CursorBench · 4.0Supporting evidence46.3%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

                Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Grok 4.7 launch comparison · Model Improvements table: CursorBench · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Reviewed 2026-09-21
                DeepSWE v1.1 · 1.1Coding71%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

                Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Grok 4.7 launch comparison · Model Improvements table: DeepSWE v1.1 · reviewed 2026-09-21

                Published configuration

                Effort: High

                lab self-report

                The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Reviewed 2026-09-21
                EEBenchSupporting evidence64%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

                Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Grok 4.7 launch comparison · Model Improvements table: EEBench · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Reviewed 2026-09-21
                GDPval-AA · 2Agentic1695
                Reported settings & source

                Grok 4.7 XHigh; Grok 4.6 High; Claude Fable 5.1 Max; GPT-6 Astra Max. GDPval launch chart.

                Launch chart rounded Elo. Original evaluator: https://artificialanalysis.ai/evaluations/gdpval-aa

                Comparison limit: Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence.

                Grok 4.7 launch comparison · Professional knowledge work: GDPval chart · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                Secondary Artificial Analysis leaderboard report; retain the original evaluator once, without treating the launch reprint as additional evidence.

                Reviewed 2026-09-21
                Harvey Legal Agent BenchmarkSupporting evidence19.6%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://www.vals.ai/benchmarks/hlab

                Comparison limit: Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

                Grok 4.7 launch comparison · Model Improvements table: Harvey Legal Agent Benchmark · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                Secondary Harvey leaderboard report; not an additional independent evaluation. Use the original Vals AI evaluation with its reviewed effort settings.

                Reviewed 2026-09-21
                HealthBench ProfessionalSupporting evidence56.7%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

                Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Grok 4.7 launch comparison · Model Improvements table: HealthBench Professional · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Reviewed 2026-09-21
                Terminal-Bench · 4.0Coding38%
                Reported settings & source

                Grok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified.

                Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High.

                Comparison limit: The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Grok 4.7 launch comparison · Model Improvements table: Terminal-Bench · reviewed 2026-09-21

                Published configuration

                Effort: XHigh

                lab self-report

                The launch table does not establish matching agent harnesses and evaluation settings. Retained as a provider claim; not a new comparable evaluation.

                Reviewed 2026-09-21
                ARC-AGI-2Hard reasoning
                DeepSWE v1.1Agentic
                GDPval-AAAgentic
                GPQA DiamondHard reasoning
                Humanity's Last ExamHard reasoning
                LiveCodeBenchCoding
                LMArena Text ArenaHuman pref
                MMLU-ProKnowledge
                OSWorld-VerifiedAgentic
                SWE-bench ProAgentic
                SWE-bench VerifiedAgentic
                Terminal-Bench 2.1Agentic