RankingThinkingBox-Bench

ThinkingBox-Bench

Data updated 6 Oct 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
2026-10-03
Display harness
ThinkingBox-Bench pass@1
Board
https://huggingface.co/blog/microsoft/thinkingbox

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

1–25 of 460 entries

ModelScoreEvidenceSource-recorded date
Claude Opus 5.5Anthropic67.16%official board2026-10-03
Claude Opus 5Anthropic66.5%official board2026-10-03
GPT-5.4OpenAI65.36%official board2026-10-03
GPT-5.6 SolOpenAI61.91%official board2026-10-03
Claude Sonnet 4.6Anthropic59.19%official board2026-10-03
GPT-6 AstraOpenAI58.31%official board2026-10-03
Kimi K3Moonshot57.37%official board2026-10-03
Qwen3.8-27BQwen51.7%official board2026-10-03
GPT-5.2OpenAI46.28%official board2026-10-03
Kimi-K2.6Moonshot37.66%official board2026-10-03
GLM-5.1Z.ai33.19%official board2026-10-03
Qwen3.6-27BQwen32.94%official board2026-10-03
Claude Opus 4.6Anthropic32.09%official board2026-10-03
o3-proOpenAI19.31%official board2026-10-03
Grok-4.3xAI14.38%official board2026-10-03
Qwen3.5-9BQwen5.65%official board2026-10-03
Mistral Large 3Mistral4.66%official board2026-10-03
Agnes 2.5 Pro AlphaAgnes AI———
Amazon Nova 2 LiteAmazon———
Amazon Nova 2 Pro PreviewAmazon———
Amazon Nova LiteAmazon———
Amazon Nova MicroAmazon———
Amazon Nova PremierAmazon———
Amazon Nova ProAmazon———
AutoGLM-Phone-9BZ.ai———