RankingAutomationBench · 1.0.6

AutomationBench · 1.0.6

Data updated 22 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
1.0.6
Display harness
Introducing GPT-6 Sol and Luna
Board
https://openai.com/index/introducing-gpt-6-sol-and-luna/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–10 of 10 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 SolOpenAI21.16%
Reported settings & source

Reported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.2116, stored as 21.16 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / low · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 LunaOpenAI1.2%
Reported settings & source

Reported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.012, stored as 1.2 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / low · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 LunaOpenAI9.4%
Reported settings & source

Reported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.094, stored as 9.4 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / medium · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 LunaOpenAI14.5%
Reported settings & source

Reported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.145, stored as 14.5 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / high · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 LunaOpenAI12.6%
Reported settings & source

Reported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.126, stored as 12.6 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / xhigh · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 LunaOpenAI20.7%
Reported settings & source

Reported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.207, stored as 20.7 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / max · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI26.94%
Reported settings & source

Reported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.2694, stored as 26.94 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / medium · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI31.2%
Reported settings & source

Reported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.312, stored as 31.2 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / high · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI33.18%
Reported settings & source

Reported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.3318, stored as 33.18 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / xhigh · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI31.96%
Reported settings & source

Reported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.3196, stored as 31.96 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / max · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22