RankingQwen 3.8-Max
Qwen 3.8-Max
50 published benchmark measures · 11 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 4 effort levels across 25 benchmark/harness combinations →
Reported effort · Max + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Max: 2 observations
- Not specified: 56 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
11 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Qwen
- Catalog status
- api-listed
- Availability
- Documented provider API; account and region restrictions may apply
- Family
- qwen3.8
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://www.alibabacloud.com/help/en/model-studio/model-pricing
- Default Capability family coverage
/badge/qwen-3.8-max.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| $OneMillion-Bench (expert score) · source release snapshot; version not specified | Supporting evidence | 52.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (ALE-CLI) · source release snapshot; version not specified | Supporting evidence | 27%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 27%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 52.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AndroidBench · source release snapshot; version not specified | Supporting evidence | 75.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Artificial Analysis Intelligence Index | Supporting evidence | 56 | AA Intelligence Index | official board | 2026-08-06 |
| Automation-Bench (Pass@1) · source release snapshot; version not specified | Agentic | 27.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (v1.0.6) · v1.0.6 | Agentic | 39.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CoWorkBench · source release snapshot; version not specified | Supporting evidence | 74.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CyberGym · source release snapshot; version not specified | Supporting evidence | 78.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSWE (v1.1) · v1.1 | Coding | 56.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE 1.1 · 1.1 | Coding | 56.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 57.5% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| ExploitBench · source release snapshot; version not specified | Supporting evidence | 28.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 14 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 26 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 73.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA | Agentic | 1630 | Artificial Analysis GDPval-AAcontributes to capability | official board | 2026-09-12 |
| GDPval-AA v2 · source release snapshot; version not specified | Agentic | 1739Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond | Hard reasoning | 92.6% | GPQA Diamond reported | lab self-report | 2026-08-03 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 92.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HealthBench · source release snapshot; version not specified | Supporting evidence | 60.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HLE · source release snapshot; version not specified | Hard reasoning | 43.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ tools · source release snapshot; version not specified | Hard reasoning | 56.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ Tools · source release snapshot; version not specified | Hard reasoning | 56.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Humanity's Last Exam | Hard reasoning | 43.6% | HLE no tools | lab self-report | 2026-08-03 |
| IFBench · source release snapshot; version not specified | Supporting evidence | 82.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 53.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LMArena Text Arena | Human pref | 1481 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| LongBench v2 · source release snapshot; version not specified | Long context | 66.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 41%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MRCR v2 256K (8-needle) · source release snapshot; version not specified | Long context | 92.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| NL2Repo · source release snapshot; version not specified | Supporting evidence | 55.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: NL2Repo · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| NL2Repo-Bench · source release snapshot; version not specified | Supporting evidence | 55.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: NL2Repo-Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OSWorld-Verified | Agentic | 86.1% | OSWorld-Verified reported | lab self-report | 2026-08-03 |
| PaperBench | Supporting evidence | 93% | PaperBench reported | lab self-report | 2026-08-03 |
| PaperBench · source release snapshot; version not specified | Supporting evidence | 93%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PLawBench · source release snapshot; version not specified | Supporting evidence | 73.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PRBench-Finance · source release snapshot; version not specified | Supporting evidence | 58.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PRBench-Legal · source release snapshot; version not specified | Supporting evidence | 57.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench (Almost Solved) · source release snapshot; version not specified | Supporting evidence | 10.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenQoderBench · source release snapshot; version not specified | Supporting evidence | 58.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenReactBench · source release snapshot; version not specified | Supporting evidence | 1724Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSVGBench · source release snapshot; version not specified | Supporting evidence | 1713Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSWEBench · source release snapshot; version not specified | Supporting evidence | 80.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SkillsBench · source release snapshot; version not specified | Supporting evidence | 70.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro | Coding | 67.7% | SWE-bench Pro reported | lab self-report | 2026-08-03 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 67.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 86.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 86.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 2.1 | Coding | 86.6% | Terminal-Bench 2.1 reported | lab self-report | 2026-08-03 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 72.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon Verified (Pass@1) · source release snapshot; version not specified | Agentic | 72.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| VulcanBench v3 | Supporting evidence | 81.2% | VulcanBench v3 bare-bones API · Report 12 · low | official board | 2026-08-04 |
| VulcanBench v3 | Supporting evidence | 71% | VulcanBench v3 bare-bones API · Report 12 · medium | official board | 2026-08-04 |
| VulcanBench v3 | Supporting evidence | 55.1% | VulcanBench v3 bare-bones API · Report 12 · xhigh | official board | 2026-08-04 |
| WideSearch · source release snapshot; version not specified | Supporting evidence | 81.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Item-F1 over4runs; Qwen-Agent for Qwen,ClaudeCode for comparators. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: WideSearch · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WorkSpaceBench · source release snapshot; version not specified | Supporting evidence | 67.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |