RankingOSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline
OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 2.0 / v2026.08.08 offline
- Display harness
- GPT-6 Astra: A new generation of intelligence
- Board
- https://openai.com/index/gpt-6-astra/
Compare published benchmark results with category weights →
Models
26–28 of 28 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6.1 SolOpenAI | 69.56%Reported settings & sourceReported reasoning effort high. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.9613; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 69.38%Reported settings & sourceReported reasoning effort xhigh. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.0466; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 71.42%Reported settings & sourceReported reasoning effort max. Offline set from v2026.08.08; partial reward, not binary full-task success or OSWorld Verified. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.2689; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: OSWorld 2.0, offline set / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |