RankingOSWorld binary · 2.0 08.08

OSWorld binary · 2.0 08.08

Data updated 12 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
2.0 08.08
Display harness
Muse Spark1.3 evaluation methodology
Board
https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Muse Spark 1.3Meta32%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Claude Opus 5Anthropic31.4%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
GPT-5.6 SolOpenAI27.3%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta26.7%
Reported settings & source

108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning xhigh

Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06