RankingMuse Spark 1.3

Muse Spark 1.3

Data updated 12 Sept 2026

11 published benchmark measures · 7 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Muse Spark 1.3 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Muse Spark 1.3 · Agentic: 64.6 · SupportedMuse Spark 1.3 · Hard reasoning: 67.0 · PreliminaryMuse Spark 1.3 · Coding: 95.0 · PreliminaryMuse Spark 1.3 · Long context: 94.4 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic4 families · 0 with independent evidence · Supported64.6

4 core families; 3 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 53.9–87.1; 0/11 scenarios unsupported. Without one publisher: 56.0–67.6; 1/15 unsupported. Smoothing check: 59.4–67.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning1 families · 1 with independent evidence · Preliminary67.0

1 core families; 22 direct opponents across 11 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GPQA: 63.6 observed win share
    artificialanalysis.ai · Source 1
Coding1 families · 0 with independent evidence · Preliminary95.0

1 core families; 1 direct opponents across 1 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 20/20 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • DeepSWE: 100.0 observed win share
    Meta · Source 1
Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      MultimodalNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Long context1 families · 0 with independent evidence · Preliminary94.4

        1 core families; 2 direct opponents across 2 labs. Needs broader benchmark and opponent support

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

        Compare 2 effort levels across 46 benchmark/harness combinations →

        Reported effort · Mixed settings

        Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

        Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

        Inspect each result and its source ↓ · Download effort evidence

        Compare capability profiles →

        Score contributions and missing evidence

        7 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.

        Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

        Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

        Model information & shareable badge
        Lab
        Meta
        Catalog status
        active
        Availability
        Documented Meta Model API and Muse Code
        Family
        Muse Spark
        Released
        2026-09-02
        Context
        License
        proprietary
        Model card
        https://research.meta.ai/blog/introducing-muse-spark-1-3
        Default Capability family coverage
        Documented-evidence family coverage/badge/muse-spark-1.3.svg
        Benchmark scores & sources

        Original results, evaluation harnesses, and evidence behind this model.

        BenchmarkBucketScoreHarnessEvidenceSource-recorded date
        Agentic IF Index · internalSupporting evidence57.8 index
        Reported settings & source

        Internal composite instruction-following evaluations; no fixed task count; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        Agentic IF Index · internalSupporting evidence55.7 index
        Reported settings & source

        Internal composite instruction-following evaluations; no fixed task count; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        AutomationBench · public v3Agentic49.6%
        Reported settings & source

        600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        AutomationBench · public v3Agentic48.3%
        Reported settings & source

        600public workflow tasks; deterministic end-state assertions;pass@1; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        contributes to capability
        lab self-reportReviewed 2026-09-06
        DeepSearchQAAgentic90.3 percent F1
        Reported settings & source

        900questions; common search backend/browser harness; answer-set F1; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        DeepSearchQAAgentic89.4 percent F1
        Reported settings & source

        900questions; common search backend/browser harness; answer-set F1; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        contributes to capability
        lab self-reportReviewed 2026-09-06
        DeepSWE · 1.1Coding75.4%
        Reported settings & source

        113tasks; Muse1.3mini-swe-agent; comparators officialDatacurve board; reasoning max

        Opus74 omitted because Google current methodology explicitly identifies that board-rounded value as incorrect; underlying precision unresolved.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        GDPval-AA · 2Agentic1754
        Reported settings & source

        Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        GDPval-AA · 2Agentic1709
        Reported settings & source

        Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        contributes to capability
        lab self-reportReviewed 2026-09-06
        GPQA DiamondHard reasoning93.535%GPQA Diamond reported

        Effort: Not specified

        contributes to capability
        independent repro2026-09-11
        JobBenchSupporting evidence64.9%
        Reported settings & source

        65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        JobBenchSupporting evidence61.2%
        Reported settings & source

        65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        MRCR · v2 256K-512KLong context98.5 percent sequence match
        Reported settings & source

        8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        MRCR · v2 256K-512KLong context97.6 percent sequence match
        Reported settings & source

        8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        contributes to capability
        lab self-reportReviewed 2026-09-06
        MRCR · v2 512K-1MLong context98.1 percent sequence match
        Reported settings & source

        8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        MRCR · v2 512K-1MLong context93.1 percent sequence match
        Reported settings & source

        8needle;100examples/band rebinned byo200k_base; no tools; sequence-matcher ratio; GPT fromOpenAIcard; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        contributes to capability
        lab self-reportReviewed 2026-09-06
        OSWorld binary · 2.0 08.08Agentic32%
        Reported settings & source

        108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

        Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        OSWorld binary · 2.0 08.08Agentic26.7%
        Reported settings & source

        108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning xhigh

        Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        No matched opponent in this evaluation unit

        Reviewed 2026-09-06
        OSWorld partial · 2.0 08.08Agentic66.9%
        Reported settings & source

        108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

        Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        contributes to capability
        lab self-reportReviewed 2026-09-06
        OSWorld partial · 2.0 08.08Agentic59%
        Reported settings & source

        108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning xhigh

        Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        No matched opponent in this evaluation unit

        Reviewed 2026-09-06
        SWE-Atlas Codebase QnASupporting evidence59.4%
        Reported settings & source

        124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        SWE-Atlas Codebase QnASupporting evidence54%
        Reported settings & source

        124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning xhigh

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        Terminal-Bench · 2.1Coding88.8%
        Reported settings & source

        89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Meta (exact harness revision not specified)

        Native harnesses differ; not a Terminus2-only comparison.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: Max

        lab self-report

        Different native coding harnesses are mixed; exact shared agent configuration is not established.

        Reviewed 2026-09-06
        Terminal-Bench · 2.1Coding89.2%
        Reported settings & source

        89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning xhigh; native harness family: Meta (exact harness revision not specified)

        Native harnesses differ; not a Terminus2-only comparison.

        Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

        Published configuration

        Effort: XHigh

        lab self-report

        Different native coding harnesses are mixed; exact shared agent configuration is not established.

        Reviewed 2026-09-06
        ARC-AGI-2Hard reasoning
        DeepSWE v1.1Agentic
        GDPval-AAAgentic
        Humanity's Last ExamHard reasoning
        LiveCodeBenchCoding
        LMArena Text ArenaHuman pref
        MMLU-ProKnowledge
        OSWorld-VerifiedAgentic
        SWE-bench ProAgentic
        SWE-bench VerifiedAgentic
        Terminal-Bench 2.1Agentic