RankingAgents' Last Exam · V1
Agents' Last Exam · V1
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- V1
- Display harness
- Introducing GPT-6 Sol and Luna
- Board
- https://openai.com/index/introducing-gpt-6-sol-and-luna/
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
1–10 of 10 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 SolOpenAI | 48.68%Reported settings & sourceReported reasoning effort low. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4868, stored as 48.68 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 36.32%Reported settings & sourceReported reasoning effort low. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.3632, stored as 36.32 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / low · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 46.84%Reported settings & sourceReported reasoning effort medium. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4684, stored as 46.84 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 43.6%Reported settings & sourceReported reasoning effort high. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.436, stored as 43.6 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 47.89%Reported settings & sourceReported reasoning effort xhigh. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.4789, stored as 47.89 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 LunaOpenAI | 50.89%Reported settings & sourceReported reasoning effort max. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5089, stored as 50.89 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Luna / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 53.06%Reported settings & sourceReported reasoning effort medium. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5306, stored as 53.06 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / medium · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 52.58%Reported settings & sourceReported reasoning effort high. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5258, stored as 52.58 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / high · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 55.39%Reported settings & sourceReported reasoning effort xhigh. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5539, stored as 55.39 percent. Not an independent board. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / xhigh · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |
| GPT-6 SolOpenAI | 56.36%Reported settings & sourceReported reasoning effort max. Agents' Last Exam V1. Long-horizon professional tasks across 55 sub-industries. OpenAI-published chart on the Introducing GPT-6 Sol and Luna page. Provider-published lab self-report. Vega-Lite chart data value 0.5636, stored as 56.36 percent. Not an independent board. Article prose rounds this max point to 56.4%. Comparison limit: Provider-published lab self-report. Shown as a labeled claim and not admitted as a matched board comparison. Introducing GPT-6 Sol and Luna · Chart: Agents' Last Exam / GPT-6 Sol / max · reviewed 2026-09-22 | Introducing GPT-6 Sol and Luna | lab self-report | 2026-09-22 |