RankingTerminal-Bench · 2.1

Terminal-Bench · 2.1

Data updated 12 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
2.1
Display harness
Muse Spark1.3 evaluation methodology
Board
https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology

Compare published benchmark results with category weights →

Models

ModelScoreHarnessEvidenceSource-recorded date
Muse Spark 1.3Meta88.8%
Reported settings & source

89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Meta (exact harness revision not specified)

Native harnesses differ; not a Terminus2-only comparison.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Claude Opus 5Anthropic89.1%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Claude Opus 5Anthropic86.7%
Reported settings & source

89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Anthropic (exact harness revision not specified)

Native harnesses differ; not a Terminus2-only comparison.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Claude Sonnet 5Anthropic80.4%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.7 FlashGoogle85.8%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Gemini 3.8 FlashGoogle89.4%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI88.8%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
GPT-5.6 SolOpenAI88.8%
Reported settings & source

89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: OpenAI (exact harness revision not specified)

Native harnesses differ; not a Terminus2-only comparison.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
GPT-5.6 TerraOpenAI87.4%
Reported settings & source

Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

Gemini3.8Flash Model Cardlab self-report2026-09-06
Muse Spark 1.2Meta82.9%
Reported settings & source

89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning xhigh; native harness family: Meta (exact harness revision not specified)

Native harnesses differ; not a Terminus2-only comparison.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06
Muse Spark 1.3Meta89.2%
Reported settings & source

89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning xhigh; native harness family: Meta (exact harness revision not specified)

Native harnesses differ; not a Terminus2-only comparison.

Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

Muse Spark1.3 evaluation methodologylab self-report2026-09-06