RankingTerminal-Bench Hard · source release snapshot; version not specified

Terminal-Bench Hard · source release snapshot; version not specified

Data updated 12 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
source release snapshot; version not specified
Display harness
Command A+ launch benchmarks
Board
https://cohere.com/blog/command-a-plus

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Command A Plus 05 2026Cohere25%
Reported settings & source

Cohere launch evaluation; North metric uses internal LLM judge.

First-party reported result; comparator results retain the source evaluation setup.

Command A+ launch benchmarks · Performance table · reviewed 2026-09-06

Command A+ launch benchmarkslab self-report2026-09-06
Command A Reasoning 08 2025Cohere3%
Reported settings & source

Cohere launch evaluation; North metric uses internal LLM judge.

First-party reported result; comparator results retain the source evaluation setup.

Command A+ launch benchmarks · Performance table · reviewed 2026-09-06

Command A+ launch benchmarkslab self-report2026-09-06