RankingQwen3.7-max

Qwen3.7-max

Data updated 12 Sept 2026

30 published benchmark measures · 8 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Qwen3.7-max capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Qwen3.7-max · Agentic: 16.0 · SupportedQwen3.7-max · Hard reasoning: 29.5 · SupportedQwen3.7-max · Coding: 22.5 · SupportedQwen3.7-max · Long context: 40.4 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic2 families · 0 with independent evidence · Supported16.0

2 core families; 3 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 15.0–19.3; 2/11 scenarios unsupported. Without one publisher: 12.5–22.1; 1/15 unsupported. Smoothing check: 10.6–24.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • AutomationBench: 0.0 observed win share
    Alibaba · Source 1
  • Toolathlon: 0.0 observed win share
    Alibaba · Source 1
Hard reasoning2 families · 0 with independent evidence · Supported29.5

2 core families; 3 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 25.1–29.9; 2/7 scenarios unsupported. Without one publisher: 26.6–32.2; 1/16 unsupported. Smoothing check: 27.4–34.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 0 with independent evidence · Supported22.5

2 core families; 3 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 22.5–26.4; 2/7 scenarios unsupported. Without one publisher: 20.9–26.5; 1/20 unsupported. Smoothing check: 19.5–28.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • DeepSWE: 0.0 observed win share
    Alibaba · Source 1
  • SWE-bench Pro: 0.0 observed win share
    Alibaba · Source 1
Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      MultimodalNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Long context2 families · 0 with independent evidence · Supported40.4

        2 core families; 3 direct opponents across 3 labs.

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: 36.6–40.4; 2/5 scenarios unsupported. Without one publisher: 40.4–40.4; 3/5 unsupported. Smoothing check: 38.5–43.6; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        • Longbench: 0.0 observed win share
          Alibaba · Source 1
        • MRCR: 33.3 observed win share
          Alibaba · Source 1

        Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

        Compare 3 effort levels across 36 benchmark/harness combinations →

        Reported effort · Not specified

        Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

        Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

        Inspect each result and its source ↓ · Download effort evidence

        Compare capability profiles →

        Score contributions and missing evidence

        8 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.

        Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

        Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

        Model information & shareable badge
        Lab
        Qwen
        Catalog status
        api-listed
        Availability
        Documented provider API; account and region restrictions may apply
        Family
        qwen3.7
        Released
        Context
        License
        proprietary
        Model card
        https://www.alibabacloud.com/help/en/model-studio/model-pricing
        Default Capability family coverage
        Documented-evidence family coverage/badge/qwen3.7-max.svg
        Benchmark scores & sources

        Original results, evaluation harnesses, and evidence behind this model.

        BenchmarkBucketScoreHarnessEvidenceSource-recorded date
        $OneMillion-Bench (expert score) · source release snapshot; version not specifiedSupporting evidence44.4%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence11.8%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        Agents' Last Exam (Pass / Score) · source release snapshot; version not specifiedSupporting evidence31.1%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        AndroidBench · source release snapshot; version not specifiedSupporting evidence56.5%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        Automation-Bench (Pass@1) · source release snapshot; version not specifiedAgentic14.2%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        CoWorkBench · source release snapshot; version not specifiedSupporting evidence64.6%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        DeepSWE 1.1 · 1.1Coding21.6%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        FrontierSWE · source release snapshot; version not specifiedSupporting evidence40.7%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        Cited comparator result; see the benchmark footnote.

        Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: FrontierSWE · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        GPQA Diamond · source release snapshot; version not specifiedHard reasoning92.4%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        HealthBench · source release snapshot; version not specifiedSupporting evidence54.5%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        HLE · source release snapshot; version not specifiedHard reasoning41.4%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        HLE w/ tools · source release snapshot; version not specifiedHard reasoning53.5%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        IFBench · source release snapshot; version not specifiedSupporting evidence79.1%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        JobBench · source release snapshot; version not specifiedSupporting evidence31.3%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        LongBench v2 · source release snapshot; version not specifiedLong context65.3%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        MLS-Bench-Lite · source release snapshot; version not specifiedSupporting evidence31.7%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        Cited comparator result; see the benchmark footnote.

        Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        MRCR v2 256K (8-needle) · source release snapshot; version not specifiedLong context86.7%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        NL2Repo-Bench · source release snapshot; version not specifiedSupporting evidence47.2%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: NL2Repo-Bench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        PaperBench · source release snapshot; version not specifiedSupporting evidence64.8%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        PLawBench · source release snapshot; version not specifiedSupporting evidence58.9%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        PRBench-Finance · source release snapshot; version not specifiedSupporting evidence46.8%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        PRBench-Legal · source release snapshot; version not specifiedSupporting evidence48.5%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        QwenQoderBench · source release snapshot; version not specifiedSupporting evidence36.8%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        QwenReactBench · source release snapshot; version not specifiedSupporting evidence1538
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        QwenSVGBench · source release snapshot; version not specifiedSupporting evidence1499
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        QwenSWEBench · source release snapshot; version not specifiedSupporting evidence63.4%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        SkillsBench · source release snapshot; version not specifiedSupporting evidence61.2%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        SWE-bench Pro · source release snapshot; version not specifiedCoding60.6%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        Terminal Bench 2.1 · 2.1Coding74.5%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses.

        Cited comparator result; see the benchmark footnote.

        Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation.

        Reviewed 2026-09-12
        Toolathlon Verified (Pass@1) · source release snapshot; version not specifiedAgentic49.7%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-12
        WideSearch · source release snapshot; version not specifiedSupporting evidence75.2%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Item-F1 over4runs; Qwen-Agent for Qwen,ClaudeCode for comparators.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: WideSearch · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        WorkSpaceBench · source release snapshot; version not specifiedSupporting evidence61.4%
        Reported settings & source

        Source benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card.

        First-party reported result.

        Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-12
        ARC-AGI-2Hard reasoning
        DeepSWE v1.1Agentic
        GDPval-AAAgentic
        GPQA DiamondHard reasoning
        Humanity's Last ExamHard reasoning
        LiveCodeBenchCoding
        LMArena Text ArenaHuman pref
        MMLU-ProKnowledge
        OSWorld-VerifiedAgentic
        SWE-bench ProAgentic
        SWE-bench VerifiedAgentic
        Terminal-Bench 2.1Agentic