RankingGPT-6.1 Sol
GPT-6.1 Sol
11 published benchmark measures · 5 benchmark families contribute across 2 task areas. 2 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High: 6 observations
- Low: 6 observations
- Max: 9 observations
- Medium: 6 observations
- XHigh: 7 observations
- Not specified: 1 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
5 contributing families across 2 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- OpenAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- GPT-6.1
- Released
- 2026-09-29
- Context
- 1,050,000 tokens
- API list price
- $2 input / $10 output per million tokens
- License
- proprietary
- Model card
- https://developers.openai.com/api/docs/models/gpt-6.1-sol
- Default Capability family coverage
/badge/gpt-6.1-sol.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AutomationBench · 1.0.6 | Agentic | 24.7%Reported settings & sourceReported reasoning effort low. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.157; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / low · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| AutomationBench · 1.0.6 | Agentic | 31.7%Reported settings & sourceReported reasoning effort medium. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1917; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| AutomationBench · 1.0.6 | Agentic | 33.2%Reported settings & sourceReported reasoning effort high. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2255; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / high · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| AutomationBench · 1.0.6 | Agentic | 35.5%Reported settings & sourceReported reasoning effort xhigh. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2508; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| AutomationBench · 1.0.6 | Agentic | 36.1%Reported settings & sourceReported reasoning effort max. End-to-end business workflows using 47 tools across six business functions. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2989; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: AutomationBench / GPT-6.1 Sol / max · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| Bug Hunt Bench | Supporting evidence | 41.9% | Bug Hunt Bench · max | official board | 2026-09-29 |
| Computer-use safety stress test · Sol 6.1 launch / safety stress test | Supporting evidence | 4.32%Reported settings & sourceReported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| DeepSWE v1.1 · v1.1 | Coding | 64.38%Reported settings & sourceReported reasoning effort low. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1714; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / low · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| DeepSWE v1.1 · v1.1 | Coding | 73.01%Reported settings & sourceReported reasoning effort medium. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.4196; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| DeepSWE v1.1 · v1.1 | Coding | 75.22%Reported settings & sourceReported reasoning effort high. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.6461; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / high · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| DeepSWE v1.1 · v1.1 | Coding | 71.9%Reported settings & sourceReported reasoning effort xhigh. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.7886; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| DeepSWE v1.1 · v1.1 | Coding | 71.9%Reported settings & sourceReported reasoning effort max. Complex software-engineering tasks in original real codebases. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.5711; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: DeepSWE / GPT-6.1 Sol / max · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations | Supporting evidence | 7.72%Reported settings & sourceReported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0452; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / low · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations | Supporting evidence | 6.29%Reported settings & sourceReported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0558; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations | Supporting evidence | 4.52%Reported settings & sourceReported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0815; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / high · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations | Supporting evidence | 4.12%Reported settings & sourceReported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations | Supporting evidence | 4.61%Reported settings & sourceReported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1301; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| Failure to disclose a broken search tool · Sol 6.1 launch / safety stress test | Supporting evidence | 2.1%Reported settings & sourceReported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29 | Published configuration | lab self-report | Reviewed 2026-09-29 |
| gdp.pdf · not specified | Agentic | 27%Reported settings & sourceReported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3341; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / low · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| gdp.pdf · not specified | Agentic | 30%Reported settings & sourceReported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3375; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| gdp.pdf · not specified | Agentic | 32%Reported settings & sourceReported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3494; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / high · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| gdp.pdf · not specified | Agentic | 31.8%Reported settings & sourceReported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3681; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| gdp.pdf · not specified | Agentic | 31%Reported settings & sourceReported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.4199; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / max · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 58.96%Reported settings & sourceReported reasoning effort low. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.4248; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / low · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 66.84%Reported settings & sourceReported reasoning effort medium. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.7675; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-29 |