Rankinggdp.pdf · not specified

gdp.pdf · not specified

Data updated 29 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
not specified
Display harness
GPT-5.6: Frontier intelligence that scales with your ambition
Board
https://openai.com/index/gpt-5-6/

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–25 of 27 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-5.6 SolOpenAI30.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Sol · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
Claude Fable 5Anthropic29.8%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Fable 5 · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
GPT-5.5OpenAI26%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.5 · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
GPT-5.6 TerraOpenAI24.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Terra · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
GPT-5.6 LunaOpenAI22.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Luna · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
Claude Opus 4.8Anthropic22.5%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Claude Opus 4.8 · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
Gemini 3.1 ProGoogle16.7%
Reported settings & source

Launch table reported configuration; per-cell reasoning effort unspecified

Provider-published result; comparator measurements are not automatically independently reproduced.

GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / Gemini 3.1 Pro Preview · reviewed 2026-09-06

GPT-5.6: Frontier intelligence that scales with your ambitionlab self-report2026-09-06
Claude Opus 5.5Anthropic25.6%
Reported settings & source

Reported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $0.76411812; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
Claude Opus 5.5Anthropic25.6%
Reported settings & source

Reported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $0.7953176399999999; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
Claude Opus 5.5Anthropic28.800000000000004%
Reported settings & source

Reported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $0.82542564; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
Claude Opus 5.5Anthropic26.6%
Reported settings & source

Reported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $0.9642728800000001; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
Claude Opus 5.5Anthropic26.2%
Reported settings & source

Reported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $1.55233548; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / Opus 5.5 w/ fallbacks / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI30.4%
Reported settings & source

Reported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.6963186; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI30.4%
Reported settings & source

Reported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.7198086; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI31%
Reported settings & source

Reported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.7948240999999998; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI32.2%
Reported settings & source

Reported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.9127230999999998; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI31%
Reported settings & source

Reported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.0758207; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Astra / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI21.8%
Reported settings & source

Reported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.33; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI25.4%
Reported settings & source

Reported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.34; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI28%
Reported settings & source

Reported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.35; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI23.8%
Reported settings & source

Reported reasoning effort xhigh. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.37; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI24.8%
Reported settings & source

Reported reasoning effort max. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.43; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI27%
Reported settings & source

Reported reasoning effort low. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3341; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI30%
Reported settings & source

Reported reasoning effort medium. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3375; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI32%
Reported settings & source

Reported reasoning effort high. Professional questions about complex PDFs across ten professional domains. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $0.3494; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: GDP.pdf / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29