RankingCommand A Plus 05 2026
Command A Plus 05 2026
16 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 1 effort levels across 19 benchmark/harness combinations →
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 16 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Cohere
- Catalog status
- active
- Availability
- Documented provider API; downloadable official weights
- Family
- Command
- Released
- —
- Context
- 128,000 tokens
- License
- Apache-2.0
- Model card
- https://docs.cohere.com/docs/models
- Default Capability family coverage
/badge/command-a-plus-05-2026.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA Intelligence Index · source release snapshot; version as labeled | Supporting evidence | 37 index pointsReported settings & sourceCohere quoted AA launch snapshot; index version unspecified. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Article prose · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AIME 2025 · source release snapshot; version as labeled | Hard reasoning | 90%Reported settings & sourceOfficial30questions x10 repeats; pass@1. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image3 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| CharXiv descriptive · source release snapshot; version as labeled | Multimodal | 88%Reported settings & sourceStandard methodology; integer rounded labels in chart. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| CharXiv Reasoning · source release snapshot; version not specified | Multimodal | 52.7%Reported settings & sourceCohere launch multimodal evaluations. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IFBench · source release snapshot; version as labeled | Supporting evidence | 74%Reported settings & sourceSingle-turn loose,prompt accuracy,294prompts x5 repeats. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MathVista · source release snapshot; version not specified | Multimodal | 80.6%Reported settings & sourceCohere launch multimodal evaluations. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU · source release snapshot; version not specified | Multimodal | 75.1%Reported settings & sourceCohere launch multimodal evaluations. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 63%Reported settings & sourceCohere launch multimodal evaluations. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MT-AIME 2025 Arabic/Japanese/Korean · source release snapshot; version as labeled | Supporting evidence | 86%Reported settings & sourceInternal Command A Translate translations; Arabic,Japanese,Korean. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image6 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| North Agentic Question Answering · source release snapshot; version as labeled | Supporting evidence | 65%Reported settings & sourceInternal North enterprise MCP cloud-file QA,LLM judge. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image4 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| North Data Analysis · source release snapshot; version as labeled | Supporting evidence | 45%Reported settings & sourceInternal North uploaded spreadsheet data-science tasks,LLM judge. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image4 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| North Memory Usage Quality · source release snapshot; version not specified | Supporting evidence | 54%Reported settings & sourceCohere launch evaluation; North metric uses internal LLM judge. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SciCode · source release snapshot; version as labeled | Coding | 38%Reported settings & source65problems/288subproblems; scientist-annotated background. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image3 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| tau2-Bench Telecom · source release snapshot; version not specified | Supporting evidence | 85%Reported settings & sourceCohere launch evaluation; North metric uses internal LLM judge. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Hard · source release snapshot; version not specified | Coding | 25%Reported settings & sourceCohere launch evaluation; North metric uses internal LLM judge. First-party reported result; comparator results retain the source evaluation setup. Command A+ launch benchmarks · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| WMT24++ 50 varieties · source release snapshot; version as labeled | Supporting evidence | 81 xCOMETxl scoreReported settings & sourcexCOMETxl average50 varieties,including internal Irish/Maltese translations and Serbian transliteration. First-party chart numeric label, visually verified. Command A+ launch benchmarks · Image6 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |