RankingOSWorld binary · 2.0 06.24

OSWorld binary · 2.0 06.24

Data updated 12 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
2.0 06.24
Display harness
Muse Spark1.3 evaluation methodology
Board
https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Muse Spark 1.2Meta17.9%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning xhigh

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06