RankingMuse Spark 1.1

Muse Spark 1.1

Data updated 12 Sept 2026

20 published benchmark measures · 13 benchmark families contribute across 6 task areas. 6 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Muse Spark 1.1 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Muse Spark 1.1 · Agentic: 44.3 · SupportedMuse Spark 1.1 · Hard reasoning: 82.5 · PreliminaryMuse Spark 1.1 · Coding: 40.6 · SupportedMuse Spark 1.1 · Human pref: 58.2 · SupportedMuse Spark 1.1 · Multimodal: 53.0 · PreliminaryMuse Spark 1.1 · Long context: 29.2 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic6 families · 1 with independent evidence · Supported44.3

6 core families; 8 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 39.3–58.1; 0/11 scenarios unsupported. Without one publisher: 43.3–46.5; 1/15 unsupported. Smoothing check: 42.1–47.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning1 families · 0 with independent evidence · Preliminary82.5

1 core families; 3 direct opponents across 3 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding3 families · 2 with independent evidence · Supported40.6

3 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 31.9–45.0; 0/7 scenarios unsupported. Without one publisher: 39.3–44.0; 0/20 unsupported. Smoothing check: 39.7–42.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported58.2

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 53.8–64.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 97.4 observed win share
    lmarena.ai · Source 1
KnowledgeNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Multimodal1 families · 0 with independent evidence · Preliminary53.0

    1 core families; 3 direct opponents across 3 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • CharXiv: 66.7 observed win share
      Meta · Source 1
    Long context1 families · 0 with independent evidence · Preliminary29.2

    1 core families; 2 direct opponents across 2 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    • MRCR: 50.0 observed win share
      Meta · Source 1

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 2 effort levels across 50 benchmark/harness combinations →

    Reported effort · Not specified

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    13 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Meta
    Catalog status
    availability-unverified
    Availability
    Documented Meta Model API; public preview at release; current version-specific serving requires authenticated verification
    Family
    Muse Spark
    Released
    2026-07-09
    Context
    1,000,000 tokens
    License
    proprietary
    Model card
    https://research.meta.ai/blog/introducing-muse-spark-meta-model-api
    Default Capability family coverage
    Documented-evidence family coverage/badge/muse-spark-1.1.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    BabyVision with tools · source release snapshot; version not specifiedSupporting evidence76.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    CharXiv Reasoning with tools · source release snapshot; version not specifiedMultimodal88.4%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSearchQA · source release snapshot; version not specifiedAgentic84.9%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1Coding53.3%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding53.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence57.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval-AA v2 Elo · source release snapshot; version not specifiedAgentic1381
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HealthBench Professional · source release snapshot; version not specifiedSupporting evidence59.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE with tools · source release snapshot; version not specifiedHard reasoning62.1%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    HLE without tools · source release snapshot; version not specifiedHard reasoning52.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    JobBench · source release snapshot; version not specifiedSupporting evidence54.7%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LMArena Text ArenaHuman pref1494LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    MCPAtlas · source release snapshot; version not specifiedAgentic88.1%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MRCR v2 1M 8-needle · source release snapshot; version not specifiedLong context54.1%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 binary without exec · 2.0Agentic14.2%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld 2.0 partial without exec · 2.0Agentic47.3%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld Verified · source release snapshot; version not specifiedAgentic80.8%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld-VerifiedAgentic80.67%OSWorld-Verified reported

    Effort: Not specified

    contributes to capability
    official board2026-07-09
    SWE-bench ProCoding61.5%SWE-bench Pro reported

    Effort: Not specified

    contributes to capability
    official board2026-07-09
    SWE-bench Pro · source release snapshot; version not specifiedCoding61.5%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 2.1 · 2.1Coding80%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Toolathlon Verified · source release snapshot; version not specifiedAgentic75.6%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    WebArena Verified · source release snapshot; version not specifiedAgentic69%
    Reported settings & source

    Meta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts.

    First-party reported result; comparator results retain the source evaluation setup.

    Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-2Hard reasoning
    GDPval-AAAgentic
    GPQA DiamondHard reasoning
    Humanity's Last ExamHard reasoning
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    SWE-bench VerifiedAgentic
    Terminal-Bench 2.1Agentic