RankingOSWorld partial · 2.0 08.08

OSWorld partial · 2.0 08.08

Data updated 12 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
2.0 08.08
Display harness
Muse Spark1.3 evaluation methodology
Board
https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Claude Opus 5Anthropic68.3%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta66.9%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
GPT-5.6 SolOpenAI62.7%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta59%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning xhigh

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06