RankingDeepSWE v1.1 · v1.1
DeepSWE v1.1 · v1.1
- Bucket
- Coding
- Unit
- percent
- Direction
- Higher is better
- Version
- v1.1
- Display harness
- GPT-6 Astra: A new generation of intelligence
- Board
- https://openai.com/index/gpt-6-astra/
Compare published benchmark results with category weights →
Models
26–38 of 38 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 SolOpenAI | 65.27%Reported settings & sourceReported reasoning effort high. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6527, stored as 65.27 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 66.59%Reported settings & sourceReported reasoning effort xhigh. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6659, stored as 66.59 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 68.81%Reported settings & sourceReported reasoning effort max. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6881, stored as 68.81 percent. Not an independent board. Article prose rounds this max point to 68.8%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 37.17%Reported settings & sourceReported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1623; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 56.64%Reported settings & sourceReported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3798; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 65.27%Reported settings & sourceReported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.6404; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 66.59%Reported settings & sourceReported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.0033; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 68.81%Reported settings & sourceReported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.7439; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 64.38%Reported settings & sourceReported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1714; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 73.01%Reported settings & sourceReported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.4196; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 75.22%Reported settings & sourceReported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.6461; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 71.9%Reported settings & sourceReported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.7886; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 71.9%Reported settings & sourceReported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.5711; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |