RankingAutomationBench · 1.0.6

AutomationBench · 1.0.6

Data updated 30 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
1.0.6
Display harness
Introducing GPT-6 Sol and Luna
Board
https://openai.com/index/introducing-gpt-6-sol-and-luna/

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

26–31 of 31 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 SolOpenAI31.96%
Reported settings & source

Reported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3406; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI24.7%
Reported settings & source

Reported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.157; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI31.7%
Reported settings & source

Reported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.1917; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI33.2%
Reported settings & source

Reported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.2255; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI35.5%
Reported settings & source

Reported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.2508; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI36.1%
Reported settings & source

Reported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.2989; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29