RankingGDPval-AA v2 (Elo) · source release snapshot; version not specified
GDPval-AA v2 (Elo) · source release snapshot; version not specified
- Bucket
- Agentic
- Unit
- elo
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- moonshotai/Kimi-K3
- Board
- https://huggingface.co/moonshotai/Kimi-K3
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 1747Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 1736Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Kimi K3Moonshot | 1686Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 1593Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 1510Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.5OpenAI | 1491Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 1741Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 1728Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| Grok 4.6xAI | 1753Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |
| Grok-4.5xAI | 1526Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Grok 4.6 launch evaluations | lab self-report | 2026-09-06 |