RankingAutoResearchExam

AutoResearchExam

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
2026-09-09
Display harness
AutoResearchExam Terminus 2
Board
https://benchmarks.bespokelabs.ai/autoresearchexam/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

201–225 of 455 entries

ModelScoreEvidenceSource-recorded date
GPT-5.5OpenAI
GPT-5.5 ProOpenAI
GPT-5.6 LunaOpenAI
GPT-5.6 TerraOpenAI
GPT-6 LunaOpenAI
GPT-6 SolOpenAI
gpt-oss-120bOpenAI
gpt-oss-20bOpenAI
Grok 4.7xAI
Grok-4.20-0309-non-reasoningxAI
Grok-4.20-0309-reasoningxAI
Grok-4.20-multi-agent-0309xAI
Grok-4.3xAI
Grok-4.5xAI
Grok-build-0.1xAI
Hunyuan A52B InstructTencent
Hunyuan-0.5B-InstructTencent
Hunyuan-1.8B-InstructTencent
Hunyuan-4B-InstructTencent
Hunyuan-7B-InstructTencent
Hunyuan-7B-Instruct-0124Tencent
Hunyuan-A13B-InstructTencent
Hy3Tencent
Hy3-previewTencent
Hy4-previewTencent