RankingAutomationBench · 1.0.6
AutomationBench · 1.0.6
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 1.0.6
- Display harness
- Introducing GPT-6 Sol and Luna
- Board
- https://openai.com/index/introducing-gpt-6-sol-and-luna/
Compare published benchmark results with category weights →
Models
26–31 of 31 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 SolOpenAI | 31.96%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3406; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 24.7%Reported settings & sourceReported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.157; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 31.7%Reported settings & sourceReported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1917; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 33.2%Reported settings & sourceReported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2255; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 35.5%Reported settings & sourceReported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2508; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 36.1%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2989; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |