RankingFrontierSWE · source release snapshot; version not specified
FrontierSWE · source release snapshot; version not specified
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- Qwen/Qwen3.8-2.4T-A95B
- Board
- https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 88.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen 3.8-MaxQwen | 73.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 70%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen3.7-maxQwen | 40.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 86.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 88.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. The source column includes fallback execution without a documented exact effort. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 66.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 66.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 67.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 67.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GLM-5.3Z.ai | 78.1%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Proximal; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: FrontierSWE · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GPT-5.5OpenAI | 64.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites FrontierSWE leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 71.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Kimi K3Moonshot | 81.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |