RankingGPT-6.1 Sol

GPT-6.1 Sol

Compare models

Data updated 30 Sept 2026

11 published benchmark measures · 5 benchmark families contribute across 2 task areas. 2 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-6.1 Sol capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-6.1 Sol · Agentic: 93.6 · PreliminaryGPT-6.1 Sol · Coding: 96.9 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic4 families · 0 with independent evidence · Preliminary93.6

4 core families; 2 direct opponents across 1 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 13/13 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoningNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Coding1 families · 0 with independent evidence · Preliminary96.9

    1 core families; 2 direct opponents across 1 labs. Needs broader benchmark and opponent support

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Human prefNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      KnowledgeNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        MultimodalNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          Long contextNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

            Reported effort · Mixed settings

            Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

            Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

            Inspect each result and its source ↓ · Download effort evidence
            Score contributions and missing evidence

            5 contributing families across 2 capabilities. Fixed reference panels do not change when the catalog expands.

            Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

            Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

            Model information & shareable badge
            Lab
            OpenAI
            Catalog status
            active
            Availability
            Public provider catalog; account and region restrictions may apply
            Family
            GPT-6.1
            Released
            2026-09-29
            Context
            1,050,000 tokens
            API list price
            $2 input / $10 output per million tokens
            Standard rate for requests up to 272,000 input tokens. Longer requests are priced at 2× input and 1.5× output for the whole request.
            License
            proprietary
            Model card
            https://developers.openai.com/api/docs/models/gpt-6.1-sol
            Default Capability family coverage
            Documented-evidence family coverage/badge/gpt-6.1-sol.svg
            Benchmark scores & sources

            Original results, evaluation harnesses, and evidence behind this model.

            BenchmarkBucketScoreHarnessEvidenceSource-recorded date
            AutomationBench · 1.0.6Agentic24.7%
            Reported settings & source

            Reported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.157; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / low · reviewed 2026-09-29

            Published configuration

            Effort: Low

            contributes to capability
            lab self-reportReviewed 2026-09-29
            AutomationBench · 1.0.6Agentic31.7%
            Reported settings & source

            Reported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.1917; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / medium · reviewed 2026-09-29

            Published configuration

            Effort: Medium

            contributes to capability
            lab self-reportReviewed 2026-09-29
            AutomationBench · 1.0.6Agentic33.2%
            Reported settings & source

            Reported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.2255; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / high · reviewed 2026-09-29

            Published configuration

            Effort: High

            contributes to capability
            lab self-reportReviewed 2026-09-29
            AutomationBench · 1.0.6Agentic35.5%
            Reported settings & source

            Reported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.2508; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

            Published configuration

            Effort: XHigh

            contributes to capability
            lab self-reportReviewed 2026-09-29
            AutomationBench · 1.0.6Agentic36.1%
            Reported settings & source

            Reported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.2989; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / max · reviewed 2026-09-29

            Published configuration

            Effort: Max

            contributes to capability
            lab self-reportReviewed 2026-09-29
            Bug Hunt BenchSupporting evidence41.9%Bug Hunt Bench · max

            Effort: Not specified

            official board

            Supporting evidence outside the reviewed capability core

            2026-09-29
            Computer-use safety stress test · Sol 6.1 launch / safety stress testSupporting evidence4.32%
            Reported settings & source

            Reported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Safety failure behavior is not a capability or intelligence score.

            Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

            Published configuration

            Effort: XHigh

            lab self-report

            Safety failure behavior is not a capability or intelligence score.

            Reviewed 2026-09-29
            DeepSWE v1.1 · v1.1Coding64.38%
            Reported settings & source

            Reported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.1714; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / low · reviewed 2026-09-29

            Published configuration

            Effort: Low

            contributes to capability
            lab self-reportReviewed 2026-09-29
            DeepSWE v1.1 · v1.1Coding73.01%
            Reported settings & source

            Reported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.4196; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / medium · reviewed 2026-09-29

            Published configuration

            Effort: Medium

            contributes to capability
            lab self-reportReviewed 2026-09-29
            DeepSWE v1.1 · v1.1Coding75.22%
            Reported settings & source

            Reported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.6461; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / high · reviewed 2026-09-29

            Published configuration

            Effort: High

            contributes to capability
            lab self-reportReviewed 2026-09-29
            DeepSWE v1.1 · v1.1Coding71.9%
            Reported settings & source

            Reported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.7886; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

            Published configuration

            Effort: XHigh

            contributes to capability
            lab self-reportReviewed 2026-09-29
            DeepSWE v1.1 · v1.1Coding71.9%
            Reported settings & source

            Reported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $1.5711; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / max · reviewed 2026-09-29

            Published configuration

            Effort: Max

            contributes to capability
            lab self-reportReviewed 2026-09-29
            Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversationsSupporting evidence7.72%
            Reported settings & source

            Reported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.0452; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / low · reviewed 2026-09-29

            Published configuration

            Effort: Low

            lab self-report

            Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Reviewed 2026-09-29
            Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversationsSupporting evidence6.29%
            Reported settings & source

            Reported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.0558; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / medium · reviewed 2026-09-29

            Published configuration

            Effort: Medium

            lab self-report

            Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Reviewed 2026-09-29
            Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversationsSupporting evidence4.52%
            Reported settings & source

            Reported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.0815; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / high · reviewed 2026-09-29

            Published configuration

            Effort: High

            lab self-report

            Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Reviewed 2026-09-29
            Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversationsSupporting evidence4.12%
            Reported settings & source

            Reported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.0999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

            Published configuration

            Effort: XHigh

            lab self-report

            Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Reviewed 2026-09-29
            Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversationsSupporting evidence4.61%
            Reported settings & source

            Reported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.1301; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29

            Published configuration

            Effort: Max

            lab self-report

            Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

            Reviewed 2026-09-29
            Failure to disclose a broken search tool · Sol 6.1 launch / safety stress testSupporting evidence2.1%
            Reported settings & source

            Reported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Comparison limit: Safety failure behavior is not a capability or intelligence score.

            Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29

            Published configuration

            Effort: Max

            lab self-report

            Safety failure behavior is not a capability or intelligence score.

            Reviewed 2026-09-29
            gdp.pdf · not specifiedAgentic27%
            Reported settings & source

            Reported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.3341; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / low · reviewed 2026-09-29

            Published configuration

            Effort: Low

            contributes to capability
            lab self-reportReviewed 2026-09-29
            gdp.pdf · not specifiedAgentic30%
            Reported settings & source

            Reported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.3375; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / medium · reviewed 2026-09-29

            Published configuration

            Effort: Medium

            contributes to capability
            lab self-reportReviewed 2026-09-29
            gdp.pdf · not specifiedAgentic32%
            Reported settings & source

            Reported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.3494; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / high · reviewed 2026-09-29

            Published configuration

            Effort: High

            contributes to capability
            lab self-reportReviewed 2026-09-29
            gdp.pdf · not specifiedAgentic31.8%
            Reported settings & source

            Reported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.3681; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

            Published configuration

            Effort: XHigh

            contributes to capability
            lab self-reportReviewed 2026-09-29
            gdp.pdf · not specifiedAgentic31%
            Reported settings & source

            Reported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.4199; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / max · reviewed 2026-09-29

            Published configuration

            Effort: Max

            contributes to capability
            lab self-reportReviewed 2026-09-29
            OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic58.96%
            Reported settings & source

            Reported reasoning effort low. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.4248; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / low · reviewed 2026-09-29

            Published configuration

            Effort: Low

            contributes to capability
            lab self-reportReviewed 2026-09-29
            OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic66.84%
            Reported settings & source

            Reported reasoning effort medium. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

            Provider-published lab self-report. Published chart cost per task $0.7675; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

            Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / medium · reviewed 2026-09-29

            Published configuration

            Effort: Medium

            contributes to capability
            lab self-reportReviewed 2026-09-29