RankingGPT-5.6 Sol
GPT-5.6 Sol
153 published benchmark measures · 25 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 7 effort levels across 72 benchmark/harness combinations →
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Best across efforts: 33 observations
- Max: 81 observations
- Not specified: 133 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
25 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- OpenAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- GPT-5.6
- Released
- —
- Context
- 1,050,000 tokens
- License
- proprietary
- Model card
- https://developers.openai.com/api/docs/models/all
- Default Capability family coverage
/badge/gpt-5.6-sol.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| $OneMillion-Bench (expert score) · source release snapshot; version not specified | Supporting evidence | 53.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: $OneMillion-Bench (expert score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA Intelligence Index · source release snapshot; version not specified | Supporting evidence | 61 index pointsReported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AA-Briefcase | Supporting evidence | 1502Reported settings & sourceArtificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase (Elo) · source release snapshot; version not specified | Supporting evidence | 1495Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-Briefcase (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase (Elo) · source release snapshot; version not specified | Supporting evidence | 1502Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AA-LCR · source release snapshot; version not specified | Long context | 73.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: AA-LCR · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agentic IF Index · internal | Supporting evidence | 60.5 indexReported settings & sourceInternal composite instruction-following evaluations; no fixed task count; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agents' Last Exam · not specified | Supporting evidence | 53.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agents' Last Exam · not specified | Supporting evidence | 52.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agents' Last Exam · source release snapshot; version not specified | Supporting evidence | 29.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites official Agents Last Exam leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Agents' Last Exam · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (ALE-CLI) · source release snapshot; version not specified | Supporting evidence | 28.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Agents' Last Exam (ALE-CLI) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 30.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Pass First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agents' Last Exam (Pass / Score) · source release snapshot; version not specified | Supporting evidence | 53.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Score First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Agents' Last Exam (Pass / Score) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AndroidBench · source release snapshot; version not specified | Supporting evidence | 74%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: AndroidBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 39.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis / APEX-Agents leaderboard; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: APEX-Agents · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 56.7%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI · 1 | Hard reasoning | 96.5%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI · 2 | Hard reasoning | 92.5%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-1 · 1 | Hard reasoning | 97.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 92.5% | ARC Prize verifiedcontributes to capability | official board | 2026-07-09 |
| ARC-AGI-2 · 2 | Hard reasoning | 92.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-3 · 3 | Hard reasoning | 7.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-3 · 3 | Hard reasoning | 7.78%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.1 · v1.1 | Supporting evidence | 80 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.4 · v1.4 | Supporting evidence | 65.1 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index | Supporting evidence | 61 | AA Intelligence Index | official board | 2026-08-12 |
| Artificial Analysis Intelligence Index v4.1 · v4.1 | Supporting evidence | 58.9 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1.1 · v4.1.1 | Supporting evidence | 60.9 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Automation-Bench (Pass@1) · source release snapshot; version not specified | Agentic | 29.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Automation-Bench (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench | Agentic | 19.6%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench · not specified | Agentic | 18.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · not specified | Agentic | 18.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · public v3 | Agentic | 46.7%Reported settings & source600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · source release snapshot; version not specified | Agentic | 29.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: AutomationBench · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench (v1.0.6) · v1.0.6 | Agentic | 45.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: AutomationBench (v1.0.6) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| BabyVision w/ python · source release snapshot; version not specified | Supporting evidence | 88.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: BabyVision w/ python · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BenchCAD · not specified | Supporting evidence | 83.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 70.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD (python tool) · not specified | Supporting evidence | 83.4%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Big Finance Bench · not specified | Supporting evidence | 53%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Difficult | Supporting evidence | 28.8%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BioMysteryBench · Human Difficult | Supporting evidence | 44.7%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Solvable | Supporting evidence | 86.1%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BioMysteryBench · Human Solvable | Supporting evidence | 79.5%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp | Agentic | 90.4% | BrowseComp reported | lab self-report | 2026-07-09 |
| BrowseComp · not specified | Agentic | 90.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · not specified | Agentic | 90.4%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · not specified | Agentic | 92.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Sol Ultra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 90.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Anthropic or OpenAI release reports; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: BrowseComp · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Capture-the-Flag Challenges · not specified | Supporting evidence | 96.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / Capture-the-Flag Challenges / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CharXiv (RQ) · source release snapshot; version not specified | Multimodal | 84.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv (RQ) · source release snapshot; version not specified | Multimodal | 89.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: CharXiv (RQ) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv Reasoning | Multimodal | 85.8%Reported settings & sourceNo tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Coronavirus-ACE2 Cell-Entry Screen · not specified | Supporting evidence | 0.43 composite scoreReported settings & sourceSystem-card reported configuration; reasoning effort unspecified Observed named-model performance; helpful-only checkpoint omitted as distinct noncatalog identity. GPT-6 Astra System Card · Section10.1.1.2.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CorpFin v2 · source release snapshot; version not specified | Supporting evidence | 64.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CorpFin v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CoWorkBench · source release snapshot; version not specified | Supporting evidence | 71.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal professional work tasks across science,finance,law,medical,productivity. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: CoWorkBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CritPt · source release snapshot; version not specified | Supporting evidence | 32.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: CritPt · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CursorBench · 3.2.0 | Supporting evidence | 67.2%Reported settings & sourceCursor production agent harness; independently measured by Cursor; max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| CursorBench v3.2 · v3.2 | Supporting evidence | 67.2%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CyberGym · source release snapshot; version not specified | Supporting evidence | 83.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: CyberGym · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSearchQA | Agentic | 93.1 percent F1Reported settings & source900questions; common search backend/browser harness; answer-set F1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · 1.1 | Coding | 72.7%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · 1.1 | Coding | 73%Reported settings & source113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · source release snapshot; version not specified | Coding | 73%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: DeepSWE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE (v1.1) · v1.1 | Coding | 72.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: DeepSWE (v1.1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE 1.1 · 1.1 | Coding | 73%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Best of Claude Code and mini-SWE-agent; Qwen best Claude Code; temp1,top_p.95,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: DeepSWE 1.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| DeepSWE v1.1 | Coding | 72.7% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 72.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 · v1.1 | Coding | 72.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 · v1.1 | Coding | 73%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · not specified | Supporting evidence | 78.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · not specified | Supporting evidence | 73.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · source release snapshot; version not specified | Supporting evidence | 76.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitBench (June-Aug 2026) · June-Aug2026 | Supporting evidence | 5.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; 20 vulnerabilities / 13 Chrome releases; 300-turn limit Provider-published result; comparator measurements are not automatically independently reproduced. Footnote14 separately reports11.5% with fewer turn-limit interruptions; main table remains5.5%. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench (June-Aug 2026) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitBench (June-Aug 2026) · June-Aug2026 | Supporting evidence | 11.5%Reported settings & sourceSimilar settings with fewer300-turn-limit interruptions; production safeguards absent Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra: A new generation of intelligence · Footnote14 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 30.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; v1 offline environment; no runtime package installation; token capped, no wall-clock cap Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 33.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; six-hour evaluation cap; alpha API latency rescaled to public API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitGym / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 24.9%Reported settings & sourceTwo-hour cap; alpha API latency rescaled; reduced safeguards Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-5.6: Frontier intelligence that scales with your ambition · Pushing the frontier on cyber and science paragraph · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 216 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 2h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ExploitGym (2h / 6h) · source release snapshot; version not specified | Supporting evidence | 293 countReported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. 6h Single-run869tasks; API time normalized by model TPS plus overhead; reported number solved; domain whitelist. First-party reported result. zai-org/GLM-5.3 · Performance table: ExploitGym (2h / 6h) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 53.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 53.8%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Extended (score) · 1.1 | Supporting evidence | 60.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Main (score) · 1.1 | Supporting evidence | 47.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode v1.1 Extended · v1.1 | Supporting evidence | 60.6%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 1-3 (v2) · v2 | Hard reasoning | 89%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 83%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 83%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierSWE · 2 | Supporting evidence | 0.32 fractionReported settings & sourceProximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| FrontierSWE · source release snapshot; version not specified | Supporting evidence | 71.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: FrontierSWE · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDP.PDF | Supporting evidence | 40%Reported settings & sourceAll-pass rate; allmodels selfcomputed byGoogle Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| gdp.pdf · not specified | Supporting evidence | 30.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | 1624 | Artificial Analysis GDPval-AAcontributes to capability | official board | 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1711Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1710Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1710Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 · source release snapshot; version not specified | Agentic | 1730Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. Externally evaluated result, attributed in the source footnote. Comparison limit: The GLM card attributes this evaluation to Artificial Analysis; it is not a new Z.ai comparison. zai-org/GLM-5.3 · Performance table: GDPval-AA v2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 · v2 | Agentic | 1748Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 (Elo) · source release snapshot; version not specified | Agentic | 1736Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: GDPval-AA v2 (Elo) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA v2 (Elo) · source release snapshot; version not specified | Agentic | 1728Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · not specified | Supporting evidence | 28.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · v13 | Supporting evidence | 32.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / GeneBench Pro / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 94.141% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 94.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · not specified | Hard reasoning | 94.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 94.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 94.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: GPQA Diamond · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GraphWalks BFS 1mil f1 · not specified | Long context | 77.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GraphWalks BFS 256k f1 · not specified | Long context | 90.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey LAB (Vals) · source release snapshot; version not specified | Supporting evidence | 2.5%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Harvey Lab-AA · source release snapshot; version not specified | Supporting evidence | 87.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Harvey Lab-AA · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Harvey Legal Agent Benchmark · source release snapshot; version not specified | Supporting evidence | 2.5%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 57 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 1764 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 55.6 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 1764 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-sol / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · source release snapshot; version not specified | Supporting evidence | 55.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HealthBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Consensus · not specified | Supporting evidence | 95.5 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 1740 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Consensus · not specified | Supporting evidence | 95.3 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 1740 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-sol / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 33.1 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 1751 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 31.1 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 1751 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-sol / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 60.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 60.5 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 3228 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-sol / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 64.1 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 3228 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-sol / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional (length-adjusted) · not specified | Supporting evidence | 60.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE · source release snapshot; version not specified | Hard reasoning | 47.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ tools · source release snapshot; version not specified | Hard reasoning | 58%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: HLE w/ tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE w/ Tools · source release snapshot; version not specified | Hard reasoning | 64.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: HLE w/ Tools · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE-Full · source release snapshot; version not specified | Hard reasoning | 44.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE-Full · source release snapshot; version not specified | Hard reasoning | 58%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: HLE-Full · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| HLE-Verified · source release snapshot; version not specified | Hard reasoning | 54.5%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| IFBench · source release snapshot; version not specified | Supporting evidence | 72.7%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: IFBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Internal Data Science Tasks · not specified | Supporting evidence | 30.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Internal Data Science Tasks / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Database Migration Tasks · not specified | Supporting evidence | 42.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Design Tasks · not specified | Supporting evidence | 47.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Internal Design Tasks / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Research Debugging Evaluation · not specified | Supporting evidence | 68.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / Internal Research Debugging Evaluation / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| JobBench | Supporting evidence | 45.4%Reported settings & source65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 45.4%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| JobBench · source release snapshot; version not specified | Supporting evidence | 45.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: JobBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| KernelGen 1P · not specified | Supporting evidence | 61.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / KernelGen 1P / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Kimi Code Bench 2.0 · 2.0 | Supporting evidence | 64.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Kimi Code Bench 2.0 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LABBench · 2 | Supporting evidence | 82.1%Reported settings & sourceSelfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Legal Research Bench · source release snapshot; version not specified | Supporting evidence | 48.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Legal Research Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| LifeSciBench · Gold v1 | Supporting evidence | 59.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / LifeSciBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LifeSciBench · not specified | Supporting evidence | 59.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1482 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| LongBench v2 · source release snapshot; version not specified | Long context | 67.1%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: LongBench v2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| LVBench · static | Multimodal | 82.1%Reported settings & sourceNo tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024 Frame budgets differ; table labels Gemini3.8static explicitly. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Management Consulting Tasks (Internal) · not specified | Supporting evidence | 43.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MathVision · source release snapshot; version not specified | Supporting evidence | 95.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MathVision · source release snapshot; version not specified | Supporting evidence | 97.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MathVision · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MCP-Atlas · source release snapshot; version not specified | Agentic | 83.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MCP-Atlas · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MCPMark-Verified · source release snapshot; version not specified | Supporting evidence | 92.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MCPMark-Verified · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MedChemBench (Internal) · not specified | Supporting evidence | 47.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MedChemBench (Internal) · not specified | Supporting evidence | 48.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 46.2%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MLS-Bench-Lite · source release snapshot; version not specified | Supporting evidence | 46.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MLS-Bench-Lite · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MMMU Pro (no tools) · not specified | Multimodal | 83%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; no tools Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU Pro (with tools) · not specified | Multimodal | 84.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (with tools) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 83%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 84.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMMU-Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| MMVU · source release snapshot; version not specified | Supporting evidence | 81.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: MMVU · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| MRCR · v2 256K-512K | Long context | 91.5 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MRCR · v2 512K-1M | Long context | 73.8 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MRCR v2 256K (8-needle) · source release snapshot; version not specified | Long context | 93.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: MRCR v2 256K (8-needle) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| NanoGPT · not specified | Supporting evidence | 9.69%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / NanoGPT / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| No-CoT math time horizon · not specified | Supporting evidence | 3.6 minutesReported settings & sourceUK AISI; single forward pass; no chain of thought Time-horizon estimate; possible contamination noted by evaluator. Higher means harder tasks solved, not slower inference. GPT-6 Astra System Card · Section9.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OfficeQA Pro · source release snapshot; version not specified | Agentic | 63.2%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OfficeQA Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OmniDocBench · source release snapshot; version not specified | Supporting evidence | 85.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OmniDocBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OpenAI MRCR v2 8-needle 256K-512K · v2 | Long context | 91.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 256K-512K Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 256K-512K · v2 | Long context | 91.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; 8-needle 256K-512K Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 512K-1M · v2 | Long context | 73.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 512K-1M Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 512K-1M · v2 | Long context | 73.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; 8-needle 512K-1M Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenScore String Quartets (1 - OMR-NED) · not specified | Supporting evidence | 0.19 1 - OMR-NEDReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / OpenScore String Quartets (1 - OMR-NED) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Organic Chemistry · 2 revised | Supporting evidence | 43.2%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OSWorld 2.0 · 2.0 | Agentic | 62.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 · 2.0 | Agentic | 62.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OSWorld 2.0 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 65.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld binary · 2.0 08.08 | Agentic | 27.3%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 08.08 | Agentic | 62.7%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 task revision unverified gpt-5.6-sol | Agentic | 62.6%Reported settings & sourcePartialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports. Comparison limit: Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld-Verified · source release snapshot; version not specified | Agentic | 83%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: OSWorld-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| PaperBench · source release snapshot; version not specified | Supporting evidence | 90.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. BasicAgent Code-Dev; Opus4.6 judge,3runs,12h each. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PaperBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PerceptionBench · source release snapshot; version not specified | Supporting evidence | 59.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: PerceptionBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PLawBench · source release snapshot; version not specified | Supporting evidence | 72.3%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PLawBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 34.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench · source release snapshot; version not specified | Supporting evidence | 36.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: PostTrainBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PostTrainBench Lite · not specified | Supporting evidence | 50.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / PostTrainBench Lite / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| PRBench-Finance · source release snapshot; version not specified | Supporting evidence | 55.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Finance · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| PRBench-Legal · source release snapshot; version not specified | Supporting evidence | 57.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: PRBench-Legal · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench · source release snapshot; version not specified | Supporting evidence | 77.6%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog or Vals AI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: ProgramBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProgramBench (Almost Solved) · source release snapshot; version not specified | Supporting evidence | 23%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: ProgramBench (Almost Solved) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProteinGym · Hard | Supporting evidence | 35.5 percent rank correlationReported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protocols · Troubleshooting | Supporting evidence | 56.4%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protocols · Understanding network-restricted | Supporting evidence | 63.9%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenQoderBench · source release snapshot; version not specified | Supporting evidence | 53.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal Qoder tasks,ClaudeCode,avg@5,6h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenQoderBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenReactBench · source release snapshot; version not specified | Supporting evidence | 1564Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN React benchmark,7categories,ClaudeCode,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenReactBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSVGBench · source release snapshot; version not specified | Supporting evidence | 1758Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal bilingual EN/CN SVG benchmark,render+multimodaljudge,BT/Elo. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSVGBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| QwenSWEBench · source release snapshot; version not specified | Supporting evidence | 73.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Internal software engineering,ClaudeCode,avg@3,8h,32768 output,temp1,256K. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: QwenSWEBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ResearchRubrics · source release snapshot; version not specified | Supporting evidence | 73.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ResearchRubrics · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| RSI Index · not specified | Supporting evidence | 57.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / RSI Index / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SaaS-Bench · source release snapshot; version not specified | Supporting evidence | 61.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SaaS-Bench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Sandbox Bench · September2026 internal | Supporting evidence | 4.5%Reported settings & source22 isolated CTF-style targets; protected-flag success metric Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.2.4 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SciCode · source release snapshot; version not specified | Coding | 56.1%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: SciCode · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ScreenSpot-Pro (no tools) · not specified | Multimodal | 76.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; no tools Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / ScreenSpot-Pro (no tools) / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SEC-Bench Pro · May2026 / public grader | Supporting evidence | 71.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; public grader; May2026 JavaScript subset Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SEC-Bench Pro · May2026 / public grader | Supporting evidence | 74.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration; reduced or absent production safeguards; public grader; May2026 JavaScript subset Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol Ultra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SEC-Bench Pro · May2026 / revised root-cause grader | Supporting evidence | 79.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; May2026 JavaScript subset,183 vulnerabilities; revised agent root-cause grader Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SHP2 Protein Function Prediction · 3 unpublished assay datasets | Supporting evidence | 0.3 mean R-squaredReported settings & sourceSystem-card reported configuration; reasoning effort unspecified Production-named model result; separate helpful-only checkpoint omitted because no exact catalog identity. GPT-6 Astra System Card · Section10.1.1.2.2 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SkillsBench · source release snapshot; version not specified | Supporting evidence | 73.5%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. v1.1 public87tasks,3runs; Anthropic ClaudeCode,OpenAI Codex,Qwen OpenCode. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SkillsBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SpreadsheetBench 2 · source release snapshot; version not specified | Supporting evidence | 32.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SpreadsheetBench 2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SRE-Bench · 262 binaries / 19 programs | Supporting evidence | 55.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SRE-Bench · 262 binaries / 19 programs | Supporting evidence | 68.7%Reported settings & sourcepass@4; four independent trials; all six objectives required; reduced production safeguards Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.2.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Atlas Codebase QnA | Supporting evidence | 53.5%Reported settings & source124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro | Coding | 64.6% | SWE-bench Pro reported | lab self-report | 2026-07-09 |
| SWE-bench Pro | Coding | 64.6%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-Bench Pro · not specified | Coding | 64.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Pro · source release snapshot; version not specified | Coding | 64.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Refined task set with problematic tasks corrected; all baselines rerun, Claude Code,temp1,top_p.95,256K context. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: SWE-bench Pro · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-Marathon · source release snapshot; version not specified | Supporting evidence | 39%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: SWE-Marathon · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-Marathon (v1.1) · v1.1 | Supporting evidence | 42.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: SWE-Marathon (v1.1) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 88.8%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. Qwen Claude Code avg@10,5h timeout,131072 output; comparators best published across harnesses. Cited comparator result; see the benchmark footnote. Comparison limit: The Qwen card cites an external leaderboard or release report for this comparator; it is not a new Qwen evaluation. Qwen/Qwen3.8-2.4T-A95B · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 2.1 · 2.1 | Coding | 88.8%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 2.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal Bench 3.0 · 3.0 | Coding | 34.6%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Terminal Bench 3.0 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench · 2.1 | Coding | 88.8%Reported settings & sourceTerminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 88.8%Reported settings & source89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: OpenAI (exact harness revision not specified) Native harnesses differ; not a Terminus2-only comparison. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 4.0 | Coding | 37.3%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over66tasks for Claude; GPT Codex CLI max from public board. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Comparison limit: System card section 8.6 cites the public Codex CLI result for Sol, while the Claude rows are internal Claude Code --bare reruns. Different agent harnesses and runs cannot form a matched comparison. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench · 4.0 | Coding | 37.3%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 | Coding | 88.8% | Terminal-Bench 2.1 reported | lab self-report | 2026-07-09 |
| Terminal-Bench 2.1 | Coding | 91.9% | Codex ultra | lab self-report | 2026-07-09 |
| Terminal-Bench 2.1 · 2.1 | Coding | 88.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 91.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Ultra, four-agent orchestration Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Sol Ultra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 88.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites GLM release blog, Artificial Analysis or OpenAI, per model; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Terminal-Bench 2.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench 4.0 · 4.0 | Coding | 37.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Science 0.1 · not specified | Hard reasoning | 22.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench v3.0 · v3.0 | Coding | 34.6%Reported settings & sourceGrok High; GPT-5.6 Sol Max; Fable 5 Max. Competitor figures from developer cards/public leaderboards as selected by xAI. First-party reported result; comparator results retain the source evaluation setup. Grok 4.6 launch evaluations · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench-Science · 0.1 | Coding | 22.4%Reported settings & source70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · not specified | Agentic | 58%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / GPT‑5.6 Sol · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon Verified · source release snapshot; version not specified | Agentic | 74.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon Verified (Pass@1) · source release snapshot; version not specified | Agentic | 74.9%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: Toolathlon Verified (Pass@1) · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon-Verified · source release snapshot; version not specified | Agentic | 74.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Toolathlon-Verified · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Video-MME (w. sub) · source release snapshot; version not specified | Supporting evidence | 89.5%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: Video-MME (w. sub) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| VulcanBench v3 | Supporting evidence | 87% | VulcanBench v3 bare-bones API · Report 07 · high | official board | 2026-07-12 |
| VulcanBench v3 | Supporting evidence | 78.3% | VulcanBench v3 bare-bones API · Report 07 · low | official board | 2026-07-12 |
| VulcanBench v3 | Supporting evidence | 82.6% | VulcanBench v3 bare-bones API · Report 07 · medium | official board | 2026-07-12 |
| WorkSpaceBench · source release snapshot; version not specified | Supporting evidence | 65.6%Reported settings & sourceSource benchmark-specific methodology; Qwen benchmark effort unspecified, GPT-5.6 Sol max per table header. Max in Qwen model names is not a reported effort setting. Comparator settings and source citations as documented in the model card. First-party reported result. Qwen/Qwen3.8-2.4T-A95B · Performance table: WorkSpaceBench · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| WorldVQA ForceAnswer · source release snapshot; version not specified | Supporting evidence | 41.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: WorldVQA ForceAnswer · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ZeroBench (pass@5) · source release snapshot; version not specified | Supporting evidence | 17%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. without tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ZeroBench (pass@5) · source release snapshot; version not specified | Supporting evidence | 35%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. with tools Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. First-party reported result. moonshotai/Kimi-K3 · Performance table: ZeroBench (pass@5) · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| τ³-Banking · source release snapshot; version not specified | Supporting evidence | 33%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Artificial Analysis; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: τ³-Banking · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |