RankingFactual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations
Factual error rate on difficult prompts · Sol 6.1 launch / user-flagged conversations
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Lower is better
- Version
- Sol 6.1 launch / user-flagged conversations
- Display harness
- Introducing GPT-6.1 Sol
- Board
- https://openai.com/index/introducing-gpt-6-1-sol/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–15 of 15 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 6.26%Reported settings & sourceReported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.2417; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 7.72%Reported settings & sourceReported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0452; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 11.42%Reported settings & sourceReported reasoning effort low. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0495; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 4.41%Reported settings & sourceReported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3103; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 3.9%Reported settings & sourceReported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.477; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 3.99%Reported settings & sourceReported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.6032; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 3.91%Reported settings & sourceReported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.7865; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Astra / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 6.86%Reported settings & sourceReported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0694; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 5.14%Reported settings & sourceReported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0994; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 4.52%Reported settings & sourceReported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1348; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 4.57%Reported settings & sourceReported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.179; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 6.29%Reported settings & sourceReported reasoning effort medium. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0558; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 4.52%Reported settings & sourceReported reasoning effort high. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0815; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 4.12%Reported settings & sourceReported reasoning effort xhigh. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.0999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 4.61%Reported settings & sourceReported reasoning effort max. Share of answers with at least one factual error on de-identified conversations where users flagged an earlier model error; deliberately difficult, not representative of typical usage. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.1301; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Private user-flagged prompt protocol needs separate review before capability scoring; retained as an exact lower-is-better sourced claim. Introducing GPT-6.1 Sol · Chart: Factual error rate on difficult prompts (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |