RankingAutoResearchExam

AutoResearchExam

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
2026-09-09
Display harness
AutoResearchExam Terminus 2
Board
https://benchmarks.bespokelabs.ai/autoresearchexam/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

326–350 of 455 entries

ModelScoreEvidenceSource-recorded date
Phi-4-reasoning-vision-15BMicrosoft
Phi-mini-MoE-instructMicrosoft
Phi-tiny-MoE-instructMicrosoft
Pixtral 12BMistral
Pixtral LargeMistral
Qianfan-VL-3BBaidu
Qianfan-VL-70BBaidu
Qianfan-VL-8BBaidu
QVQ-maxQwen
QVQ-plusQwen
Qwen-coder-plusQwen
Qwen-coder-turboQwen
Qwen-flashQwen
Qwen-flash-characterQwen
Qwen-longQwen
Qwen-math-plusQwen
Qwen-math-turboQwen
Qwen-maxQwen
Qwen-omni-turboQwen
Qwen-plusQwen
Qwen-plus-characterQwen
Qwen-plus-character-jaQwen
Qwen-turboQwen
Qwen-vl-maxQwen
Qwen-vl-plusQwen