RankingFinance Agent v2 · source release snapshot; version not specified
Finance Agent v2 · source release snapshot; version not specified
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- moonshotai/Kimi-K3
- Board
- https://huggingface.co/moonshotai/Kimi-K3
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Fable 5Anthropic | 56.3%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Kimi K3Moonshot | 54.4%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 53.9%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 53.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GPT-5.5OpenAI | 51.8%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 49.7%Reported settings & sourceTable headers: Kimi K3, GPT-5.6 Sol, Opus 4.8 and GLM-5.2 max; GPT-5.5 xhigh; Fable 5 max with fallbacks. Benchmark-specific footnotes override header settings; cited results are not new Moonshot evaluations. Benchmark-specific tools, harness and budget documented under Evaluation Details. Kimi temp1; single-step top_p.95,agentic top_p1; vision tools=Python,3runs except ZeroBench5runs. Cited result; see benchmark-specific Evaluation Details. Comparison limit: Kimi K3 Evaluation Details cites Vals AI; retain as published context, not a new Moonshot comparison. moonshotai/Kimi-K3 · Performance table: Finance Agent v2 · reviewed 2026-09-12 | moonshotai/Kimi-K3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 53.9%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Claude Opus 5Anthropic | 58.6%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| Claude Sonnet 5Anthropic | 53.9%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 43%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Gemini 3.7 FlashGoogle | 59%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| Gemini 3.8 FlashGoogle | 61.4%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 51.8%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 53.8%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| GPT-5.6 TerraOpenAI | 54.4%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Gemini 3.8 Flash launch performance | lab self-report | 2026-09-06 |
| Muse Spark 1.1Meta | 57.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |