RankingOSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline
OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 2.0 / v2026.08.08 offline
- Display harness
- GPT-6 Astra: A new generation of intelligence
- Board
- https://openai.com/index/gpt-6-astra/
Compare published benchmark results with category weights →
Models
1–13 of 13 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 72.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑6 Astra · reviewed 2026-09-06 | GPT-6 Astra: A new generation of intelligence | lab self-report | 2026-09-06 |
| Claude Opus 5Anthropic | 70.2%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / Claude Opus 5 · reviewed 2026-09-06 | GPT-6 Astra: A new generation of intelligence | lab self-report | 2026-09-06 |
| GPT-5.6 SolOpenAI | 65.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑5.6 Sol · reviewed 2026-09-06 | GPT-6 Astra: A new generation of intelligence | lab self-report | 2026-09-06 |
| GPT-6 LunaOpenAI | 8.26%Reported settings & sourceReported reasoning effort low. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.0826, stored as 8.26 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 31.54%Reported settings & sourceReported reasoning effort medium. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3154, stored as 31.54 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 41.43%Reported settings & sourceReported reasoning effort high. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4143, stored as 41.43 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 46.7%Reported settings & sourceReported reasoning effort xhigh. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.467, stored as 46.7 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 52.68%Reported settings & sourceReported reasoning effort max. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5268, stored as 52.68 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Luna / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 43.9%Reported settings & sourceReported reasoning effort low. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.439, stored as 43.9 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 54%Reported settings & sourceReported reasoning effort medium. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.54, stored as 54 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 58.29%Reported settings & sourceReported reasoning effort high. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5829, stored as 58.29 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 60.54%Reported settings & sourceReported reasoning effort xhigh. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6054, stored as 60.54 percent. Not an independent board. Article prose rounds this xhigh point to 60.5%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 64.43%Reported settings & sourceReported reasoning effort max. Partial reward on the offline set from the v2026.08.08 release. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.6443, stored as 64.43 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: OSWorld 2.0, offline set / GPT-6 Sol / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |