Rankinggdp.pdf · not specified
gdp.pdf · not specified
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- not specified
- Display harness
- GPT-5.6: Frontier intelligence that scales with your ambition
- Board
- https://openai.com/index/gpt-5-6/
Compare published benchmark results with category weights →
Models
1–25 of 27 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-5.6 SolOpenAI | 30.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Sol · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| Claude Fable 5Anthropic | 29.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Fable 5 · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 26%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.5 · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| GPT-5.6 TerraOpenAI | 24.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Terra · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| GPT-5.6 LunaOpenAI | 22.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Luna · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| Claude Opus 4.8Anthropic | 22.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Opus 4.8 · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 16.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Gemini 3.1 Pro Preview · reviewed 2026-09-06 | GPT-5.6: Frontier intelligence that scales with your ambition | lab self-report | 2026-09-06 |
| Claude Opus 5.5Anthropic | 25.6%Reported settings & sourceReported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.76411812; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 25.6%Reported settings & sourceReported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.7953176399999999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 28.800000000000004%Reported settings & sourceReported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.82542564; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 26.6%Reported settings & sourceReported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $0.9642728800000001; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| Claude Opus 5.5Anthropic | 26.2%Reported settings & sourceReported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $1.55233548; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 30.4%Reported settings & sourceReported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.6963186; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 30.4%Reported settings & sourceReported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.7198086; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 31%Reported settings & sourceReported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.7948240999999998; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 32.2%Reported settings & sourceReported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.9127230999999998; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 31%Reported settings & sourceReported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.0758207; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 21.8%Reported settings & sourceReported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.33; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 25.4%Reported settings & sourceReported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.34; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 28%Reported settings & sourceReported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.35; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 23.8%Reported settings & sourceReported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.37; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 24.8%Reported settings & sourceReported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.43; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 27%Reported settings & sourceReported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3341; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 30%Reported settings & sourceReported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3375; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 32%Reported settings & sourceReported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $0.3494; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |