RankingTerminal Bench 2.1 · 2.1
Terminal Bench 2.1 · 2.1
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- 2.1
- Display harness
- Qwen/Qwen3.8-2.4T-A95B
- Board
- https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-5.6 SolOpenAI | 88.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen 3.8-MaxQwen | 86.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 84.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. The source states that Fable 5 results may involve fallback execution; an exact configuration is not established. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 84.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Qwen3.7-maxQwen | 74.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Qwen/Qwen3.8-2.4T-A95B | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 88%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. Comparison limit: The source labels this column Fable-5 (w/ fallback); it does not establish standalone Fable 5 performance. Retained as published evidence, excluded from comparable ranking inputs. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 88%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. Comparison limit: The source column includes fallback execution without a documented exact effort. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 85%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 85%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| DeepSeek V4-Pro 0813DeepSeek | 87.9%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| DeepSeek V4-Pro 0813DeepSeek | 87.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| DeepSeek-V4-Flash-0731DeepSeek | 82.7%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| GLM-5.1Z.ai | 59.3%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 59.3 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| GLM-5.2Z.ai | 81%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 81%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GLM-5.3Z.ai | 88.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 88.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Kimi K3Moonshot | 88.3%Reported settings & sourceDeepSeek updated 0813/0731 release evaluation table; tool/harness settings per model card. First-party reported result. deepseek-ai/DeepSeek-V4-Pro-0813 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | deepseek-ai/DeepSeek-V4-Pro-0813 | lab self-report | 2026-09-12 |
| Kimi K3Moonshot | 88.3%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Kimi-K2.6Moonshot | 67.2%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 67.2 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| MiniMax-M2.7MiniMax | 55.5%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 55.5 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| NVIDIA Nemotron 3 UltraNVIDIA | 56.4%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 56.4 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |
| Qwen 3.8-MaxQwen | 86.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Qwen3.5-397B-A17BQwen | 49.9%Reported settings & sourceNVIDIA evaluation harness/settings per benchmark in model card; comparator scores are NVIDIA-reported under that protocol. First-party reported result. Original cell: 49.9 The linked public reproduction recipe names Terminal Bench 2.0, while this card labels 2.1. Recipe settings are not transferred across that unresolved version mismatch. nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-06 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | lab self-report | 2026-09-06 |