RankingDeepSWE v1.1 · v1.1

DeepSWE v1.1 · v1.1

Data updated 30 Sept 2026

Bucket
Coding
Unit
percent
Direction
Higher is better
Version
v1.1
Display harness
GPT-6 Astra: A new generation of intelligence
Board
https://openai.com/index/gpt-6-astra/

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

26–38 of 38 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 SolOpenAI65.27%
Reported settings & source

Reported reasoning effort high. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.6527, stored as 65.27 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / high · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI66.59%
Reported settings & source

Reported reasoning effort xhigh. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.6659, stored as 66.59 percent. Not an independent board.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / xhigh · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI68.81%
Reported settings & source

Reported reasoning effort max. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

Provider-published lab self-report. Vega-Lite chart data value 0.6881, stored as 68.81 percent. Not an independent board. Article prose rounds this max point to 68.8%.

Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / max · reviewed 2026-09-22

Introducing GPT-6 Sol and Lunalab self-report2026-09-22
GPT-6 SolOpenAI37.17%
Reported settings & source

Reported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.1623; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI56.64%
Reported settings & source

Reported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3798; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI65.27%
Reported settings & source

Reported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.6404; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI66.59%
Reported settings & source

Reported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.0033; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI68.81%
Reported settings & source

Reported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.7439; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI64.38%
Reported settings & source

Reported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.1714; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI73.01%
Reported settings & source

Reported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.4196; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI75.22%
Reported settings & source

Reported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.6461; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI71.9%
Reported settings & source

Reported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.7886; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI71.9%
Reported settings & source

Reported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.5711; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29