RankingAutoResearchExam

AutoResearchExam

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
2026-09-09
Display harness
AutoResearchExam Terminus 2
Board
https://benchmarks.bespokelabs.ai/autoresearchexam/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

76–100 of 455 entries

ModelScoreEvidenceSource-recorded date
DeepSeek-R1-Distill-Qwen-7BDeepSeek
DeepSeek-V3.2DeepSeek
DeepSeek-V3.2-SpecialeDeepSeek
DeepSeek-V4-Flash-0731DeepSeek
DeepSeek-V4-Flash-Vision-ExpDeepSeek
Devstral 2Mistral
Devstral Medium 1.0Mistral
Devstral Small 1.1Mistral
Devstral Small 2Mistral
DiffusionGemma 26B A4BGoogle
Doubao Seed CodeByteDance
ERNIE-4.5-0.3B-PTBaidu
ERNIE-4.5-21B-A3B-PTBaidu
ERNIE-4.5-21B-A3B-ThinkingBaidu
ERNIE-4.5-300B-A47B-PTBaidu
ERNIE-4.5-turbo-128kBaidu
ERNIE-4.5-turbo-20260402Baidu
ERNIE-4.5-turbo-32kBaidu
ERNIE-4.5-turbo-vlBaidu
ERNIE-4.5-turbo-vl-32kBaidu
ERNIE-4.5-VL-28B-A3B-PTBaidu
ERNIE-4.5-VL-28B-A3B-ThinkingBaidu
ERNIE-4.5-VL-424B-A47B-PTBaidu
ERNIE-5.0Baidu
ERNIE-5.0-thinking-expBaidu