RankingAutoResearchExam

AutoResearchExam

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
2026-09-09
Display harness
AutoResearchExam Terminus 2
Board
https://benchmarks.bespokelabs.ai/autoresearchexam/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

1–25 of 455 entries

ModelScoreEvidenceSource-recorded date
Claude Fable 5.1Anthropic0.602 official board2026-09-09
GPT-6 AstraOpenAI0.6 official board2026-09-09
Claude Opus 5Anthropic0.579 official board2026-09-09
GPT-5.6 SolOpenAI0.513 official board2026-09-09
Grok 4.6xAI0.493 official board2026-09-09
Gemini 3.8 FlashGoogle0.484 official board2026-09-09
Kimi K3Moonshot0.435 official board2026-09-09
Qwen 3.8-MaxQwen0.423 official board2026-09-09
Muse Spark 1.3Meta0.379 official board2026-09-09
Agnes 2.5 Pro AlphaAgnes AI
Amazon Nova 2 LiteAmazon
Amazon Nova 2 Pro PreviewAmazon
Amazon Nova LiteAmazon
Amazon Nova MicroAmazon
Amazon Nova PremierAmazon
Amazon Nova ProAmazon
AutoGLM-Phone-9BZ.ai
AutoGLM-Phone-9B-MultilingualZ.ai
aya-101Cohere
aya-23-35BCohere
aya-23-8BCohere
aya-expanse-8bCohere
aya-vision-8bCohere
C4Ai Aya Expanse 32BCohere
C4Ai Aya Vision 32BCohere