RankingComputer-use safety stress test · Sol 6.1 launch / safety stress test
Computer-use safety stress test · Sol 6.1 launch / safety stress test
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Lower is better
- Version
- Sol 6.1 launch / safety stress test
- Display harness
- Introducing GPT-6.1 Sol
- Board
- https://openai.com/index/introducing-gpt-6-1-sol/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 2.4%Reported settings & sourceReported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6 Astra / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 4.32%Reported settings & sourceReported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 LunaOpenAI | 13.7%Reported settings & sourceReported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6 Luna / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 17.39%Reported settings & sourceReported reasoning effort xhigh. Unintended outcomes in deliberately adversarial computer- and browser-use workplace tasks; the updated harder safety subset, not OSWorld task success. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Computer-use safety stress test (lower is better) / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |