RankingAutomationBench · 1.0.6
AutomationBench · 1.0.6
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 1.0.6
- Display harness
- Introducing GPT-6 Sol and Luna
- Board
- https://openai.com/index/introducing-gpt-6-sol-and-luna/
Compare published benchmark results with category weights →
Models
1–25 of 31 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 SolOpenAI | 21.16%Reported settings & sourceReported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.2116, stored as 21.16 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 1.2%Reported settings & sourceReported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.012, stored as 1.2 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| Claude Fable 5.1Anthropic | 31.4%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Fable 5.1 w/ Opus 5 fallback. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $2.45; source-specific cost does not establish consensus workload efficiency. The launch caption says fallback costs are omitted and fallbacks occurred on about 40% of tasks. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Fable 5.1 w/ Opus 5 fallback / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 24.2%Reported settings & sourceReported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.51; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Opus 5.5 w/ fallbacks / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 29.53%Reported settings & sourceReported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.65; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Opus 5.5 w/ fallbacks / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 33.03%Reported settings & sourceReported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.71; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Opus 5.5 w/ fallbacks / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 35.77%Reported settings & sourceReported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.89; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Opus 5.5 w/ fallbacks / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 42.47%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $1.44; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: AutomationBench / Opus 5.5 w/ fallbacks / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 30.3%Reported settings & sourceReported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.08; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Astra / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 34.1%Reported settings & sourceReported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.27; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Astra / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 37.1%Reported settings & sourceReported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.44; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Astra / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 39%Reported settings & sourceReported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.5; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Astra / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 41.4%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.73; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Astra / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 LunaOpenAI | 9.4%Reported settings & sourceReported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.094, stored as 9.4 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 14.5%Reported settings & sourceReported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.145, stored as 14.5 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 12.6%Reported settings & sourceReported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.126, stored as 12.6 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 20.7%Reported settings & sourceReported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.207, stored as 20.7 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 26.94%Reported settings & sourceReported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.2694, stored as 26.94 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 31.2%Reported settings & sourceReported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.312, stored as 31.2 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 33.18%Reported settings & sourceReported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3318, stored as 33.18 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 31.96%Reported settings & sourceReported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3196, stored as 31.96 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 21.16%Reported settings & sourceReported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1861; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 26.94%Reported settings & sourceReported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2098; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 31.2%Reported settings & sourceReported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2368; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 33.18%Reported settings & sourceReported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2746; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |