RankingOSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline

OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline

Data updated 30 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
2.0 / v2026.08.08 offline
Display harness
GPT-6 Astra: A new generation of intelligence
Board
https://openai.com/index/gpt-6-astra/

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

26–28 of 28 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6.1 SolOpenAI69.56%
Reported settings & source

Reported reasoning effort high. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.9613; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI69.38%
Reported settings & source

Reported reasoning effort xhigh. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.0466; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI71.42%
Reported settings & source

Reported reasoning effort max. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.2689; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29