RankingGPT-6 Luna

GPT-6 Luna

Compare models

Data updated 22 Sept 2026

5 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Compare with

Choose a second model and its effort to see both profiles and benchmark differences below.

Compare with

Choose a configuration to compare. The overview below combines evidence across settings and has no model-wide rank.

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-6 Luna capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Mixed settings

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                OpenAI
                Catalog status
                active
                Availability
                Public provider catalog; account and region restrictions may apply
                Family
                GPT-6
                Released
                2026-09-22
                Context
                1,050,000 tokens
                API list price
                $0.10 input / $0.50 output per million tokens
                Standard rate for requests up to 272,000 input tokens. Longer requests are priced at 2× input and 1.5× output for the whole request.
                License
                proprietary
                Model card
                https://developers.openai.com/api/docs/models/gpt-6-luna
                Default Capability family coverage
                Documented-evidence family coverage/badge/gpt-6-luna.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                Agents' Last Exam · V1Supporting evidence36.32%
                Reported settings & source

                Reported reasoning effort low. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.3632, stored as 36.32 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / low · reviewed 2026-09-22

                Published configuration

                Effort: Low

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Agents' Last Exam · V1Supporting evidence46.84%
                Reported settings & source

                Reported reasoning effort medium. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.4684, stored as 46.84 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / medium · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Agents' Last Exam · V1Supporting evidence43.6%
                Reported settings & source

                Reported reasoning effort high. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.436, stored as 43.6 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / high · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Agents' Last Exam · V1Supporting evidence47.89%
                Reported settings & source

                Reported reasoning effort xhigh. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.4789, stored as 47.89 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / xhigh · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                Agents' Last Exam · V1Supporting evidence50.89%
                Reported settings & source

                Reported reasoning effort max. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.5089, stored as 50.89 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / max · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBench · 1.0.6Agentic1.2%
                Reported settings & source

                Reported reasoning effort low. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.012, stored as 1.2 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / low · reviewed 2026-09-22

                Published configuration

                Effort: Low

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBench · 1.0.6Agentic9.4%
                Reported settings & source

                Reported reasoning effort medium. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.094, stored as 9.4 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / medium · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBench · 1.0.6Agentic14.5%
                Reported settings & source

                Reported reasoning effort high. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.145, stored as 14.5 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / high · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBench · 1.0.6Agentic12.6%
                Reported settings & source

                Reported reasoning effort xhigh. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.126, stored as 12.6 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / xhigh · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                AutomationBench · 1.0.6Agentic20.7%
                Reported settings & source

                Reported reasoning effort max. AutomationBench 1.0.6. End-to-end workflows across 47 tools in sales, marketing, operations, support, finance, and HR. The page says the Claude Fable 5.1 cost point omits Opus 5 fallbacks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.207, stored as 20.7 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: AutomationBench / GPT-6 Luna / max · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE v1.1 · v1.1Coding2.43%
                Reported settings & source

                Reported reasoning effort low. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.0243, stored as 2.43 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Luna / low · reviewed 2026-09-22

                Published configuration

                Effort: Low

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE v1.1 · v1.1Coding44.47%
                Reported settings & source

                Reported reasoning effort medium. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.4447, stored as 44.47 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Luna / medium · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE v1.1 · v1.1Coding59.29%
                Reported settings & source

                Reported reasoning effort high. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.5929, stored as 59.29 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Luna / high · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE v1.1 · v1.1Coding61.28%
                Reported settings & source

                Reported reasoning effort xhigh. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.6128, stored as 61.28 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Luna / xhigh · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                DeepSWE v1.1 · v1.1Coding66.59%
                Reported settings & source

                Reported reasoning effort max. DeepSWE v1.1. Original long-horizon software-engineering tasks. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.6659, stored as 66.59 percent. Not an independent board. Article prose rounds this max point to 66.6%.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: DeepSWE / GPT-6 Luna / max · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode 1.1 Main (score) · 1.1 MainSupporting evidence25.66%
                Reported settings & source

                Reported reasoning effort low. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.2566, stored as 25.66 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Luna / low · reviewed 2026-09-22

                Published configuration

                Effort: Low

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode 1.1 Main (score) · 1.1 MainSupporting evidence35.53%
                Reported settings & source

                Reported reasoning effort medium. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.3553, stored as 35.53 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Luna / medium · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode 1.1 Main (score) · 1.1 MainSupporting evidence37.26%
                Reported settings & source

                Reported reasoning effort high. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.3726, stored as 37.26 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Luna / high · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode 1.1 Main (score) · 1.1 MainSupporting evidence37.1%
                Reported settings & source

                Reported reasoning effort xhigh. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.371, stored as 37.1 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Luna / xhigh · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                FrontierCode 1.1 Main (score) · 1.1 MainSupporting evidence42.42%
                Reported settings & source

                Reported reasoning effort max. FrontierCode 1.1 Main, as identified in the surrounding article. The chart title is FrontierCode. Graded on correctness and mergeability. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.4242, stored as 42.42 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: FrontierCode / GPT-6 Luna / max · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic8.26%
                Reported settings & source

                Reported reasoning effort low. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.0826, stored as 8.26 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / low · reviewed 2026-09-22

                Published configuration

                Effort: Low

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic31.54%
                Reported settings & source

                Reported reasoning effort medium. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.3154, stored as 31.54 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / medium · reviewed 2026-09-22

                Published configuration

                Effort: Medium

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic41.43%
                Reported settings & source

                Reported reasoning effort high. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.4143, stored as 41.43 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / high · reviewed 2026-09-22

                Published configuration

                Effort: High

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic46.7%
                Reported settings & source

                Reported reasoning effort xhigh. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.467, stored as 46.7 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / xhigh · reviewed 2026-09-22

                Published configuration

                Effort: XHigh

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22
                OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic52.68%
                Reported settings & source

                Reported reasoning effort max. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page.

                Provider-published lab self-report. Vega-Lite chart data value 0.5268, stored as 52.68 percent. Not an independent board.

                Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / max · reviewed 2026-09-22

                Published configuration

                Effort: Max

                lab self-report

                Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison.

                Reviewed 2026-09-22