RankingAutoResearchExam

AutoResearchExam

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
2026-09-09
Display harness
AutoResearchExam Terminus 2
Board
https://benchmarks.bespokelabs.ai/autoresearchexam/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

51–75 of 455 entries

ModelScoreEvidenceSource-recorded date
cogvlm2-llama3-chat-19BZ.ai
cogvlm2-llama3-chinese-chat-19BZ.ai
cogvlm2-video-llama3-chatZ.ai
CommandCohere
Command A 03 2025Cohere
Command A Plus 05 2026Cohere
Command A Reasoning 08 2025Cohere
Command A Translate 08 2025Cohere
Command A Vision 07 2025Cohere
Command LightCohere
Command R 03 2024Cohere
Command R 08 2024Cohere
Command R Plus 04 2024Cohere
Command R Plus 08 2024Cohere
Command R7B 12 2024Cohere
DeepSeek V4 Flash 0424DeepSeek
DeepSeek V4-Pro 0424DeepSeek
DeepSeek V4-Pro 0813DeepSeek
DeepSeek-R1-0528DeepSeek
DeepSeek-R1-0528-Qwen3-8BDeepSeek
DeepSeek-R1-Distill-Llama-70BDeepSeek
DeepSeek-R1-Distill-Llama-8BDeepSeek
DeepSeek-R1-Distill-Qwen-1.5BDeepSeek
DeepSeek-R1-Distill-Qwen-14BDeepSeek
DeepSeek-R1-Distill-Qwen-32BDeepSeek