Evidence coverage
2666 published observations across 285 measures. Capability scoring groups compatible evidence into 43 benchmark families.
A published result contributes only when a matching reference-panel result is available. Missing evidence stays unknown. Coverage describes breadth, separately from performance.
Coverage across our labs
| Lab | Models with results | Published measures | Launch sources |
|---|---|---|---|
| OpenAI | 31 / 37 | 209 | 16 |
| Anthropic | 11 / 11 | 217 | 17 |
| xAI | 3 / 7 | 18 | 1 |
| 21 / 34 | 109 | 8 | |
| Meta | 11 / 14 | 33 | 2 |
| Mistral | 16 / 27 | 15 | 1 |
| Moonshot | 3 / 7 | 119 | 7 |
| Qwen | 28 / 81 | 91 | 4 |
| DeepSeek | 4 / 15 | 28 | 2 |
| NVIDIA | 2 / 10 | 41 | 1 |
| Z.ai | 17 / 52 | 93 | 7 |
| MiniMax | 6 / 13 | 64 | 2 |
| Cohere | 10 / 26 | 18 | 1 |
| Amazon | 5 / 6 | 19 | 1 |
| Microsoft | 9 / 24 | 19 | 1 |
| Xiaomi | 3 / 11 | 3 | 0 |
| StepFun | 2 / 4 | 2 | 0 |
| LG AI Research | 5 / 13 | 3 | 0 |
| Upstage | 4 / 8 | 2 | 0 |
| ByteDance | 1 / 16 | 1 | 0 |
| Tencent | 1 / 13 | 2 | 0 |
| Baidu | 1 / 20 | 1 | 0 |
| OpenBMB | 1 / 1 | 1 | 0 |
| Agnes AI | 1 / 1 | 1 | 0 |
Models and contributing families
Raw published counts include all source types. Contributing and independent family counts follow the selected evidence policy.
7 models match “Moonshot”. Clear search
| Model | Published measures | Contributing families | With independent support | Explore |
|---|---|---|---|---|
| Kimi-K2.6 | 66 | 19 | 1 | Published results |
| Kimi K3 | 67 | 15 | 5 | Published results |
| Kimi-K2.7-Code | 1 | 1 | 1 | Published results |
| Kimi-Dev-72B | 0 | 0 | 0 | Published results |
| Kimi-Linear-48B-A3B-Instruct | 0 | 0 | 0 | Published results |
| Kimi-VL-A3B-Instruct | 0 | 0 | 0 | Published results |
| Kimi-VL-A3B-Thinking-2506 | 0 | 0 | 0 | Published results |
Family coverage
| Benchmark family | Capability | Models contributing | With independent support |
|---|---|---|---|
| OSWorld | Agentic | 19 | 6 |
| Webarena Verified | Agentic | 4 | 0 |
| Browsecomp | Agentic | 22 | 0 |
| Toolathlon | Agentic | 15 | 0 |
| MCP-Atlas | Agentic | 19 | 0 |
| AutomationBench | Agentic | 19 | 0 |
| DeepSearchQA | Agentic | 10 | 0 |
| GDPval | Agentic | 60 | 41 |
| Bfcl V4 | Agentic | 12 | 0 |
| Apex Agents | Agentic | 10 | 0 |
| OfficeQA | Agentic | 11 | 0 |
| Humanity’s Last Exam | Hard reasoning | 31 | 8 |
| GPQA | Hard reasoning | 48 | 23 |
| ARC-AGI | Hard reasoning | 48 | 48 |
| Aime | Hard reasoning | 22 | 0 |
| FrontierMath | Hard reasoning | 10 | 0 |
| Imo Answerbench | Hard reasoning | 5 | 0 |
| Terminal-Bench Science | Hard reasoning | 5 | 0 |
| SWE-bench Verified | Coding | 38 | 19 |
| SWE-bench Pro | Coding | 20 | 4 |
| SWE-bench Multilingual | Coding | 9 | 0 |
| DeepSWE | Coding | 31 | 28 |
| LiveCodeBench | Coding | 25 | 9 |
| Terminal-Bench | Coding | 41 | 2 |
| Scicode | Coding | 7 | 0 |
| Human preference (Arena) | Human pref | 115 | 115 |
| MMLU-Pro | Knowledge | 68 | 52 |
| Gmmlu | Knowledge | 3 | 0 |
| Milu | Knowledge | 3 | 0 |
| Simpleqa Verified | Knowledge | 2 | 0 |
| Mmmu | Multimodal | 2 | 0 |
| Mmmu Pro | Multimodal | 24 | 0 |
| CharXiv | Multimodal | 16 | 0 |
| Mathvista | Multimodal | 2 | 0 |
| Chartography | Multimodal | 3 | 0 |
| Lvbench | Multimodal | 6 | 0 |
| ScreenSpot | Multimodal | 12 | 0 |
| Ocrbench V2 Average Accuracy | Multimodal | 12 | 0 |
| MRCR | Long context | 12 | 0 |
| Longbench | Long context | 8 | 0 |
| Ruler 1M | Long context | 2 | 0 |
| Graphwalks | Long context | 7 | 0 |
| Aa Lcr | Long context | 5 | 0 |
Model evidence matrix
| Model | OSWorld | Webarena Verified | Browsecomp | Toolathlon | MCP-Atlas | AutomationBench | DeepSearchQA | GDPval | Bfcl V4 | Apex Agents | OfficeQA | Humanity’s Last Exam | GPQA | ARC-AGI | Aime | FrontierMath | Imo Answerbench | Terminal-Bench Science | SWE-bench Verified | SWE-bench Pro | SWE-bench Multilingual | DeepSWE | LiveCodeBench | Terminal-Bench | Scicode | Human preference (Arena) | MMLU-Pro | Gmmlu | Milu | Simpleqa Verified | Mmmu | Mmmu Pro | CharXiv | Mathvista | Chartography | Lvbench | ScreenSpot | Ocrbench V2 Average Accuracy | MRCR | Longbench | Ruler 1M | Graphwalks | Aa Lcr |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Kimi-K2.6 | 20.0 | 78.6 | 28.6 | 44.6 | 100.0 | 80.0 | 100.0 | 100.0 | 58.3 | 79.2 | 50.0 | 100.0 | 88.9 | 100.0 | 83.3 | 62.5 | 30.0 | 50.0 | 100.0 | ||||||||||||||||||||||||
| Kimi K3 | 62.5 | 93.3 | 80.0 | 86.1 | 100.0 | 85.0 | 60.0 | 45.8 | 66.8 | 63.8 | 79.4 | 77.1 | 94.7 | 62.5 | 75.0 | ||||||||||||||||||||||||||||
| Kimi-K2.7-Code | 7.4 | ||||||||||||||||||||||||||||||||||||||||||
| Kimi-Dev-72B | |||||||||||||||||||||||||||||||||||||||||||
| Kimi-Linear-48B-A3B-Instruct | |||||||||||||||||||||||||||||||||||||||||||
| Kimi-VL-A3B-Instruct | |||||||||||||||||||||||||||||||||||||||||||
| Kimi-VL-A3B-Thinking-2506 |
Family definitions and scoring rules · Compare capabilities · Compare published evidence