Evidence coverage
2666 published observations across 285 measures. Capability scoring groups compatible evidence into 43 benchmark families.
A published result contributes only when a matching reference-panel result is available. Missing evidence stays unknown. Coverage describes breadth, separately from performance.
Coverage across our labs
| Lab | Models with results | Published measures | Launch sources |
|---|---|---|---|
| OpenAI | 31 / 37 | 209 | 16 |
| Anthropic | 11 / 11 | 217 | 17 |
| xAI | 3 / 7 | 18 | 1 |
| 21 / 34 | 109 | 8 | |
| Meta | 11 / 14 | 33 | 2 |
| Mistral | 16 / 27 | 15 | 1 |
| Moonshot | 3 / 7 | 119 | 7 |
| Qwen | 28 / 81 | 91 | 4 |
| DeepSeek | 4 / 15 | 28 | 2 |
| NVIDIA | 2 / 10 | 41 | 1 |
| Z.ai | 17 / 52 | 93 | 7 |
| MiniMax | 6 / 13 | 64 | 2 |
| Cohere | 10 / 26 | 18 | 1 |
| Amazon | 5 / 6 | 19 | 1 |
| Microsoft | 9 / 24 | 19 | 1 |
| Xiaomi | 3 / 11 | 3 | 0 |
| StepFun | 2 / 4 | 2 | 0 |
| LG AI Research | 5 / 13 | 3 | 0 |
| Upstage | 4 / 8 | 2 | 0 |
| ByteDance | 1 / 16 | 1 | 0 |
| Tencent | 1 / 13 | 2 | 0 |
| Baidu | 1 / 20 | 1 | 0 |
| OpenBMB | 1 / 1 | 1 | 0 |
| Agnes AI | 1 / 1 | 1 | 0 |
Models and contributing families
Raw published counts include all source types. Contributing and independent family counts follow the selected evidence policy.
34 models match “Google”. Clear search
| Model | Published measures | Contributing families | With independent support | Explore |
|---|---|---|---|---|
| Gemma 3 270M IT | 0 | 0 | 0 | Published results |
| Gemma 3n E2B IT | 0 | 0 | 0 | Published results |
| Gemma 4 12B IT | 0 | 0 | 0 | Published results |
| Gemma 4 26B-A4B IT | 0 | 0 | 0 | Published results |
| Gemma 4 31B IT | 0 | 0 | 0 | Published results |
| Gemma 4 E2B IT | 0 | 0 | 0 | Published results |
| Gemma 4 E4B IT | 0 | 0 | 0 | Published results |
| RecurrentGemma 2B IT | 0 | 0 | 0 | Published results |
| RecurrentGemma 9B IT | 0 | 0 | 0 | Published results |
Family coverage
| Benchmark family | Capability | Models contributing | With independent support |
|---|---|---|---|
| OSWorld | Agentic | 19 | 6 |
| Webarena Verified | Agentic | 4 | 0 |
| Browsecomp | Agentic | 22 | 0 |
| Toolathlon | Agentic | 15 | 0 |
| MCP-Atlas | Agentic | 19 | 0 |
| AutomationBench | Agentic | 19 | 0 |
| DeepSearchQA | Agentic | 10 | 0 |
| GDPval | Agentic | 60 | 41 |
| Bfcl V4 | Agentic | 12 | 0 |
| Apex Agents | Agentic | 10 | 0 |
| OfficeQA | Agentic | 11 | 0 |
| Humanity’s Last Exam | Hard reasoning | 31 | 8 |
| GPQA | Hard reasoning | 48 | 23 |
| ARC-AGI | Hard reasoning | 48 | 48 |
| Aime | Hard reasoning | 22 | 0 |
| FrontierMath | Hard reasoning | 10 | 0 |
| Imo Answerbench | Hard reasoning | 5 | 0 |
| Terminal-Bench Science | Hard reasoning | 5 | 0 |
| SWE-bench Verified | Coding | 38 | 19 |
| SWE-bench Pro | Coding | 20 | 4 |
| SWE-bench Multilingual | Coding | 9 | 0 |
| DeepSWE | Coding | 31 | 28 |
| LiveCodeBench | Coding | 25 | 9 |
| Terminal-Bench | Coding | 41 | 2 |
| Scicode | Coding | 7 | 0 |
| Human preference (Arena) | Human pref | 115 | 115 |
| MMLU-Pro | Knowledge | 68 | 52 |
| Gmmlu | Knowledge | 3 | 0 |
| Milu | Knowledge | 3 | 0 |
| Simpleqa Verified | Knowledge | 2 | 0 |
| Mmmu | Multimodal | 2 | 0 |
| Mmmu Pro | Multimodal | 24 | 0 |
| CharXiv | Multimodal | 16 | 0 |
| Mathvista | Multimodal | 2 | 0 |
| Chartography | Multimodal | 3 | 0 |
| Lvbench | Multimodal | 6 | 0 |
| ScreenSpot | Multimodal | 12 | 0 |
| Ocrbench V2 Average Accuracy | Multimodal | 12 | 0 |
| MRCR | Long context | 12 | 0 |
| Longbench | Long context | 8 | 0 |
| Ruler 1M | Long context | 2 | 0 |
| Graphwalks | Long context | 7 | 0 |
| Aa Lcr | Long context | 5 | 0 |
Model evidence matrix
| Model | OSWorld | Webarena Verified | Browsecomp | Toolathlon | MCP-Atlas | AutomationBench | DeepSearchQA | GDPval | Bfcl V4 | Apex Agents | OfficeQA | Humanity’s Last Exam | GPQA | ARC-AGI | Aime | FrontierMath | Imo Answerbench | Terminal-Bench Science | SWE-bench Verified | SWE-bench Pro | SWE-bench Multilingual | DeepSWE | LiveCodeBench | Terminal-Bench | Scicode | Human preference (Arena) | MMLU-Pro | Gmmlu | Milu | Simpleqa Verified | Mmmu | Mmmu Pro | CharXiv | Mathvista | Chartography | Lvbench | ScreenSpot | Ocrbench V2 Average Accuracy | MRCR | Longbench | Ruler 1M | Graphwalks | Aa Lcr |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma 3 270M IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 3n E2B IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 4 12B IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 4 26B-A4B IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 4 31B IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 4 E2B IT | |||||||||||||||||||||||||||||||||||||||||||
| Gemma 4 E4B IT | |||||||||||||||||||||||||||||||||||||||||||
| RecurrentGemma 2B IT | |||||||||||||||||||||||||||||||||||||||||||
| RecurrentGemma 9B IT |
Family definitions and scoring rules · Compare capabilities · Compare published evidence