RankingGPT-6 Sol
GPT-6 Sol
5 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 5 observations
- Low: 5 observations
- Max: 5 observations
- Medium: 5 observations
- XHigh: 5 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- OpenAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- GPT-6
- Released
- 2026-09-22
- Context
- 1,050,000 tokens
- API list price
- $2 input / $10 output per million tokens
- License
- proprietary
- Model card
- https://developers.openai.com/api/docs/models/gpt-6-sol
- Default Capability family coverage
/badge/gpt-6-sol.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agents' Last Exam · V1 | Supporting evidence | 48.68%Reported settings & sourceReported reasoning effort low. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4868, stored as 48.68 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / low · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Agents' Last Exam · V1 | Supporting evidence | 53.06%Reported settings & sourceReported reasoning effort medium. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5306, stored as 53.06 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / medium · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Agents' Last Exam · V1 | Supporting evidence | 52.58%Reported settings & sourceReported reasoning effort high. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5258, stored as 52.58 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / high · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Agents' Last Exam · V1 | Supporting evidence | 55.39%Reported settings & sourceReported reasoning effort xhigh. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5539, stored as 55.39 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| Agents' Last Exam · V1 | Supporting evidence | 56.36%Reported settings & sourceReported reasoning effort max. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5636, stored as 56.36 percent. Not an independent board. Article prose rounds this max point to 56.4%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / max · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench · 1.0.6 | Agentic | 21.16%Reported settings & sourceReported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.2116, stored as 21.16 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / low · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench · 1.0.6 | Agentic | 26.94%Reported settings & sourceReported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.2694, stored as 26.94 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / medium · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench · 1.0.6 | Agentic | 31.2%Reported settings & sourceReported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.312, stored as 31.2 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / high · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench · 1.0.6 | Agentic | 33.18%Reported settings & sourceReported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3318, stored as 33.18 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| AutomationBench · 1.0.6 | Agentic | 31.96%Reported settings & sourceReported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3196, stored as 31.96 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Sol / max · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE v1.1 · v1.1 | Coding | 37.17%Reported settings & sourceReported reasoning effort low. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3717, stored as 37.17 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / low · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE v1.1 · v1.1 | Coding | 56.64%Reported settings & sourceReported reasoning effort medium. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5664, stored as 56.64 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / medium · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE v1.1 · v1.1 | Coding | 65.27%Reported settings & sourceReported reasoning effort high. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6527, stored as 65.27 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / high · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE v1.1 · v1.1 | Coding | 66.59%Reported settings & sourceReported reasoning effort xhigh. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6659, stored as 66.59 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| DeepSWE v1.1 · v1.1 | Coding | 68.81%Reported settings & sourceReported reasoning effort max. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6881, stored as 68.81 percent. Not an independent board. Article prose rounds this max point to 68.8%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Sol / max · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode 1.1 Main (score) · 1.1 Main | Supporting evidence | 37.28%Reported settings & sourceReported reasoning effort low. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3728, stored as 37.28 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Sol / low · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode 1.1 Main (score) · 1.1 Main | Supporting evidence | 45.92%Reported settings & sourceReported reasoning effort medium. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4592, stored as 45.92 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Sol / medium · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode 1.1 Main (score) · 1.1 Main | Supporting evidence | 47.7%Reported settings & sourceReported reasoning effort high. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.477, stored as 47.7 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Sol / high · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode 1.1 Main (score) · 1.1 Main | Supporting evidence | 48.45%Reported settings & sourceReported reasoning effort xhigh. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4845, stored as 48.45 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| FrontierCode 1.1 Main (score) · 1.1 Main | Supporting evidence | 49.27%Reported settings & sourceReported reasoning effort max. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4927, stored as 49.27 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Sol / max · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 43.9%Reported settings & sourceReported reasoning effort low. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.439, stored as 43.9 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / low · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 54%Reported settings & sourceReported reasoning effort medium. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.54, stored as 54 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / medium · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 58.29%Reported settings & sourceReported reasoning effort high. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5829, stored as 58.29 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / high · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 60.54%Reported settings & sourceReported reasoning effort xhigh. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6054, stored as 60.54 percent. Not an independent board. Article prose rounds this xhigh point to 60.5%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 64.43%Reported settings & sourceReported reasoning effort max. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6443, stored as 64.43 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / max · reviewed 2026-09-22 | Published configuration | lab self-report | Reviewed 2026-09-22 |