RankingBug Hunt Bench
Bug Hunt Bench
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- 2026-09-22
- Display harness
- Bug Hunt Bench · max
- Board
- https://bughunt.productcompass.pm/?preset=featured
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
201–225 of 455 entries
Other harnesses
These runs use other harnesses. Capability scoring matches configurations separately; model details identify which source records contribute.
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5.1Anthropic | 31.43% | Bug Hunt Bench · high | official board | 2026-09-01 |
| Claude Opus 5Anthropic | 20% | Bug Hunt Bench · high | official board | 2026-07-26 |
| Gemini 3.8 FlashGoogle | 17.14% | Bug Hunt Bench · high | official board | 2026-09-15 |
| GPT-5.6 SolOpenAI | 32.38% | Bug Hunt Bench · high | official board | 2026-07-31 |
| GPT-6 AstraOpenAI | 40.95% | Bug Hunt Bench · xhigh | official board | 2026-09-04 |
| Grok 4.6xAI | 27.33% | Bug Hunt Bench · xhigh | official board | 2026-09-14 |
| Grok 4.7xAI | 27.43% | Bug Hunt Bench · xhigh | official board | 2026-09-21 |
| Kimi K3Moonshot | 20% | Bug Hunt Bench · default | official board | 2026-07-26 |
| Muse Spark 1.3Meta | 19.33% | Bug Hunt Bench · xhigh | official board | 2026-09-14 |
| Qwen3.8-27BQwen | 14.29% | Bug Hunt Bench · xhigh | official board | 2026-09-15 |