RankingAA Briefcase · 1.1
AA Briefcase · 1.1
- Bucket
- Supporting evidence
- Unit
- elo
- Direction
- Higher is better
- Version
- 1.1
- Display harness
- Grok 4.7 launch comparison
- Board
- https://x.ai/news/grok-4-7
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 1678Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons. Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.7xAI | 1657Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons. Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| Grok 4.6xAI | 1546Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons. Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |
| GPT-5.6 SolOpenAI | 1487Reported settings & sourceGrok 4.7 XHigh (DeepSWE: High); Grok 4.6 High; GPT-5.6 Sol Max; Claude Fable 5.1 Max. Launch comparison; per-model agent harness, evaluation budget and rerun provenance are not specified. Official launch table. Effort follows the column header except the Grok 4.7 DeepSWE asterisk, which explicitly means High. Original evaluator: https://artificialanalysis.ai/evaluations/aa-briefcase Comparison limit: Secondary Artificial Analysis leaderboard report; not an additional independent evaluation. Use the original evaluator for reviewed exact-configuration comparisons. Grok 4.7 launch comparison · Model Improvements table: AA Briefcase · reviewed 2026-09-21 | Grok 4.7 launch comparison | lab self-report | 2026-09-21 |