RankingBug Hunt Bench

Bug Hunt Bench

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
2026-09-22
Display harness
Bug Hunt Bench · max
Board
https://bughunt.productcompass.pm/?preset=featured

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

451–455 of 455 entries

ModelScoreEvidenceSource-recorded date
Voxtral MiniMistral
Voxtral SmallMistral
WeDLM-7B-InstructTencent
WeDLM-8B-InstructTencent
Youtu-LLM-2BTencent

Other harnesses

These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.

ModelScoreHarnessEvidenceSource-recorded date
Claude Fable 5.1Anthropic31.43%Bug Hunt Bench · highofficial board2026-09-01
Claude Opus 5Anthropic20%Bug Hunt Bench · highofficial board2026-07-26
Gemini 3.8 FlashGoogle17.14%Bug Hunt Bench · highofficial board2026-09-15
GPT-5.6 SolOpenAI32.38%Bug Hunt Bench · highofficial board2026-07-31
GPT-6 AstraOpenAI40.95%Bug Hunt Bench · xhighofficial board2026-09-04
Grok 4.6xAI27.33%Bug Hunt Bench · xhighofficial board2026-09-14
Grok 4.7xAI27.43%Bug Hunt Bench · xhighofficial board2026-09-21
Kimi K3Moonshot20%Bug Hunt Bench · defaultofficial board2026-07-26
Muse Spark 1.3Meta19.33%Bug Hunt Bench · xhighofficial board2026-09-14
Qwen3.8-27BQwen14.29%Bug Hunt Bench · xhighofficial board2026-09-15