RankingCommand A Reasoning 08 2025

Command A Reasoning 08 2025

Data updated 12 Sept 2026

10 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Command A Reasoning 08 2025 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoning1 families · 0 with independent evidence · PreliminaryUnknown

    1 core families; 1 direct opponents across 1 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Aime: 0.0 observed win share
      Cohere · Source 1
    Coding2 families · 0 with independent evidence · PreliminaryUnknown

    2 core families; 1 direct opponents across 1 labs. No connected comparison to the complete reference panel

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 20/20 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • Scicode: 0.0 observed win share
      Cohere · Source 1
    • Terminal-Bench: 0.0 observed win share
      Cohere · Source 1
    Human prefNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      KnowledgeNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        MultimodalNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          Long contextNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

            Compare available effort levels across 0 benchmark/harness combinations →

            Reported effort · Not specified

            Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

            Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

            Inspect each result and its source ↓ · Download effort evidence

            Compare capability profiles →

            Score contributions and missing evidence

            0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

            Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

            Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

            Model information & shareable badge
            Lab
            Cohere
            Catalog status
            active
            Availability
            Documented provider API; downloadable official weights
            Family
            Command
            Released
            Context
            256,000 tokens
            License
            CC-BY-NC-4.0
            Model card
            https://docs.cohere.com/docs/models
            Default Capability family coverage
            Documented-evidence family coverage/badge/command-a-reasoning-08-2025.svg
            Benchmark scores & sources

            Original results, evaluation harnesses, and evidence behind this model.

            BenchmarkBucketScoreHarnessEvidenceSource-recorded date
            AIME 2025 · source release snapshot; version as labeledHard reasoning57%
            Reported settings & source

            Official30questions x10 repeats; pass@1.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image3 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            contributes to capability
            lab self-reportReviewed 2026-09-06
            IFBench · source release snapshot; version as labeledSupporting evidence36%
            Reported settings & source

            Single-turn loose,prompt accuracy,294prompts x5 repeats.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image3 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            MT-AIME 2025 Arabic/Japanese/Korean · source release snapshot; version as labeledSupporting evidence53%
            Reported settings & source

            Internal Command A Translate translations; Arabic,Japanese,Korean.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image6 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            North Agentic Question Answering · source release snapshot; version as labeledSupporting evidence45%
            Reported settings & source

            Internal North enterprise MCP cloud-file QA,LLM judge.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image4 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            North Data Analysis · source release snapshot; version as labeledSupporting evidence13%
            Reported settings & source

            Internal North uploaded spreadsheet data-science tasks,LLM judge.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image4 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            North Memory Usage Quality · source release snapshot; version not specifiedSupporting evidence39%
            Reported settings & source

            Cohere launch evaluation; North metric uses internal LLM judge.

            First-party reported result; comparator results retain the source evaluation setup.

            Command A+ launch benchmarks · Performance table · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            SciCode · source release snapshot; version as labeledCoding30%
            Reported settings & source

            65problems/288subproblems; scientist-annotated background.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image3 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            contributes to capability
            lab self-reportReviewed 2026-09-06
            tau2-Bench Telecom · source release snapshot; version not specifiedSupporting evidence37%
            Reported settings & source

            Cohere launch evaluation; North metric uses internal LLM judge.

            First-party reported result; comparator results retain the source evaluation setup.

            Command A+ launch benchmarks · Performance table · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            Terminal-Bench Hard · source release snapshot; version not specifiedCoding3%
            Reported settings & source

            Cohere launch evaluation; North metric uses internal LLM judge.

            First-party reported result; comparator results retain the source evaluation setup.

            Command A+ launch benchmarks · Performance table · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            contributes to capability
            lab self-reportReviewed 2026-09-06
            WMT24++ 50 varieties · source release snapshot; version as labeledSupporting evidence73 xCOMETxl score
            Reported settings & source

            xCOMETxl average50 varieties,including internal Irish/Maltese translations and Serbian transliteration.

            First-party chart numeric label, visually verified.

            Command A+ launch benchmarks · Image6 · reviewed 2026-09-06

            Published configuration

            Effort: Not specified

            lab self-report

            Supporting evidence outside the reviewed capability core

            Reviewed 2026-09-06
            ARC-AGI-2Hard reasoning
            DeepSWE v1.1Agentic
            GDPval-AAAgentic
            GPQA DiamondHard reasoning
            Humanity's Last ExamHard reasoning
            LiveCodeBenchCoding
            LMArena Text ArenaHuman pref
            MMLU-ProKnowledge
            OSWorld-VerifiedAgentic
            SWE-bench ProAgentic
            SWE-bench VerifiedAgentic
            Terminal-Bench 2.1Agentic