RankingFailure to disclose a broken search tool · Sol 6.1 launch / safety stress test

Failure to disclose a broken search tool · Sol 6.1 launch / safety stress test

Data updated 29 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Lower is better
Version
Sol 6.1 launch / safety stress test
Display harness
Introducing GPT-6.1 Sol
Board
https://openai.com/index/introducing-gpt-6-1-sol/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–4 of 4 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 AstraOpenAI1.5%
Reported settings & source

Reported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Safety failure behavior is not a capability or intelligence score.

Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6 Astra / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI2.1%
Reported settings & source

Reported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Safety failure behavior is not a capability or intelligence score.

Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI4.92%
Reported settings & source

Reported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Safety failure behavior is not a capability or intelligence score.

Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 LunaOpenAI28.67%
Reported settings & source

Reported reasoning effort max. Adversarial test of whether agents disclose a broken search tool. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Safety stress outcomes are retained for inspection and excluded from capability scoring. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Safety failure behavior is not a capability or intelligence score.

Introducing GPT-6.1 Sol · Chart: Failure to disclose a broken search tool (lower is better) / GPT-6 Luna / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29