RankingFactual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations

Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations

Data updated 29 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Lower is better
Version
Sol 6.1 launch / user-flagged conversations
Display harness
Introducing GPT-6.1 Sol
Board
https://openai.com/index/introducing-gpt-6-1-sol/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–15 of 15 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 AstraOpenAI6.26%
Reported settings & source

Reported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.2417; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI7.72%
Reported settings & source

Reported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0452; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI11.42%
Reported settings & source

Reported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0495; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI4.41%
Reported settings & source

Reported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3103; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI3.9%
Reported settings & source

Reported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.477; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI3.99%
Reported settings & source

Reported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.6032; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI3.91%
Reported settings & source

Reported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.7865; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI6.86%
Reported settings & source

Reported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0694; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI5.14%
Reported settings & source

Reported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0994; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI4.52%
Reported settings & source

Reported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.1348; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI4.57%
Reported settings & source

Reported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.179; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI6.29%
Reported settings & source

Reported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0558; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI4.52%
Reported settings & source

Reported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0815; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI4.12%
Reported settings & source

Reported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.0999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI4.61%
Reported settings & source

Reported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.1301; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim.

Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29