RankingGrok-4.5
Grok-4.5
13 published benchmark measures · 5 benchmark families contribute across 3 task areas. 3 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 3 effort levels across 42 benchmark/harness combinations →
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 15 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
5 contributing families across 3 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- xAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Grok 4
- Released
- —
- Context
- 500,000 tokens
- License
- proprietary
- Model card
- https://docs.x.ai/developers/pricing
- Default Capability family coverage
/badge/grok-4.5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA Intelligence Index · source release snapshot; version not specified | Supporting evidence | 56 index pointsReported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AA-Briefcase (Elo) · source release snapshot; version not specified | Supporting evidence | 1313Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 47.1%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| APEX-SWE · source release snapshot; version not specified | Supporting evidence | 53.6%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CursorBench v3.2 · v3.2 | Supporting evidence | 66.7%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 53.8% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 54%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierCode v1.1 Extended · v1.1 | Supporting evidence | 56.6%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 (Elo) · source release snapshot; version not specified | Agentic | 1526Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey LAB (Vals) · source release snapshot; version not specified | Supporting evidence | 12.9%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1469 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| Terminal-Bench v3.0 · v3.0 | Coding | 15.7%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| VulcanBench v3 | Supporting evidence | 91.3% | VulcanBench v3 bare-bones API · Report 07 · high | official board | 2026-07-12 |
| VulcanBench v3 | Supporting evidence | 82.6% | VulcanBench v3 bare-bones API · Report 07 · low | official board | 2026-07-12 |
| VulcanBench v3 | Supporting evidence | 91.3% | VulcanBench v3 bare-bones API · Report 07 · medium | official board | 2026-07-12 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |