RankingReviewer bypass attempts · Sol 6.1 launch / safety stress test
Reviewer bypass attempts · Sol 6.1 launch / safety stress test
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Lower is better
- Version
- Sol 6.1 launch / safety stress test
- Display harness
- Introducing GPT-6.1 Sol
- Board
- https://openai.com/index/introducing-gpt-6-1-sol/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–4 of 4 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 0%Reported settings & sourceReported reasoning effort max. Attempts to work around an automated safety reviewer blocking an action. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Reviewer bypass attempts (lower is better) / GPT-6 Astra / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 0%Reported settings & sourceReported reasoning effort max. Attempts to work around an automated safety reviewer blocking an action. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Reviewer bypass attempts (lower is better) / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 0%Reported settings & sourceReported reasoning effort max. Attempts to work around an automated safety reviewer blocking an action. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Reviewer bypass attempts (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 LunaOpenAI | 0.26%Reported settings & sourceReported reasoning effort max. Attempts to work around an automated safety reviewer blocking an action. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Safety failure behavior is not a capability or intelligence score. Introducing GPT-6.1 Sol · Chart: Reviewer bypass attempts (lower is better) / GPT-6 Luna / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |