RankingMuse Spark 1.3
Muse Spark 1.3
11 published benchmark measures · 7 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 2 effort levels across 46 benchmark/harness combinations →
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Max: 12 observations
- XHigh: 11 observations
- Not specified: 1 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
7 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Meta
- Catalog status
- active
- Availability
- Documented Meta Model API and Muse Code
- Family
- Muse Spark
- Released
- 2026-09-02
- Context
- —
- License
- proprietary
- Model card
- https://research.meta.ai/blog/introducing-muse-spark-1-3
- Default Capability family coverage
/badge/muse-spark-1.3.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agentic IF Index · internal | Supporting evidence | 57.8 indexReported settings & sourceInternal composite instruction-following evaluations; no fixed task count; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agentic IF Index · internal | Supporting evidence | 55.7 indexReported settings & sourceInternal composite instruction-following evaluations; no fixed task count; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · public v3 | Agentic | 49.6%Reported settings & source600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · public v3 | Agentic | 48.3%Reported settings & source600public workflow tasks; deterministic end-state assertions;pass@1; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA | Agentic | 90.3 percent F1Reported settings & source900questions; common search backend/browser harness; answer-set F1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSearchQA | Agentic | 89.4 percent F1Reported settings & source900questions; common search backend/browser harness; answer-set F1; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · 1.1 | Coding | 75.4%Reported settings & source113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1754Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1709Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 93.535% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| JobBench | Supporting evidence | 64.9%Reported settings & source65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| JobBench | Supporting evidence | 61.2%Reported settings & source65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MRCR · v2 256K-512K | Long context | 98.5 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MRCR · v2 256K-512K | Long context | 97.6 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MRCR · v2 512K-1M | Long context | 98.1 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MRCR · v2 512K-1M | Long context | 93.1 percent sequence matchReported settings & source8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld binary · 2.0 08.08 | Agentic | 32%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld binary · 2.0 08.08 | Agentic | 26.7%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning xhigh Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 08.08 | Agentic | 66.9%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 08.08 | Agentic | 59%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning xhigh Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Atlas Codebase QnA | Supporting evidence | 59.4%Reported settings & source124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Atlas Codebase QnA | Supporting evidence | 54%Reported settings & source124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning xhigh Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 88.8%Reported settings & source89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Meta (exact harness revision not specified) Native harnesses differ; not a Terminus2-only comparison. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 89.2%Reported settings & source89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning xhigh; native harness family: Meta (exact harness revision not specified) Native harnesses differ; not a Terminus2-only comparison. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |