RankingTerminal-Bench Science 0.1 · 0.1

Terminal-Bench Science 0.1 · 0.1

Data updated 29 Sept 2026

Bucket
Agentic
Unit
percent
Direction
Higher is better
Version
0.1
Display harness
GPT-6 Astra: A new generation of intelligence
Board
https://openai.com/index/gpt-6-astra/

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–17 of 17 entries

ModelScoreHarnessEvidenceSource-recorded date
GPT-6 AstraOpenAI61.1%
Reported settings & source

Lower-cost setting; exact effort not specified

Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

GPT-6 Astra: A new generation of intelligence · Terminal-Bench Science chart caption · reviewed 2026-09-06

GPT-6 Astra: A new generation of intelligencelab self-report2026-09-06
Claude Opus 5.5Anthropic63.33%
Reported settings & source

Reported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone.

Provider-published lab self-report. Published chart cost per task $23.2108; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Comparison limit: Combined fallback setup is not the named base model at one effort setting.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / Opus 5.5 w/ fallbacks / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI55.43%
Reported settings & source

Reported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $11.4061; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI57.43%
Reported settings & source

Reported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $12.3383; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI62%
Reported settings & source

Reported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $14.9534; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI60.86%
Reported settings & source

Reported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $15.7554; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 AstraOpenAI68.1%
Reported settings & source

Reported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $23.7974; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI9.17%
Reported settings & source

Reported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.9986; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI14.49%
Reported settings & source

Reported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $4.4068; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI14.61%
Reported settings & source

Reported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $4.6325; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI25.29%
Reported settings & source

Reported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $6.7698; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6 SolOpenAI27.59%
Reported settings & source

Reported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $12.1803; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI43.71%
Reported settings & source

Reported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $1.7918; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / low · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI47.56%
Reported settings & source

Reported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.3386; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / medium · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI51.14%
Reported settings & source

Reported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.7594; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / high · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI53.71%
Reported settings & source

Reported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $2.8922; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / xhigh · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29
GPT-6.1 SolOpenAI57.02%
Reported settings & source

Reported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration.

Provider-published lab self-report. Published chart cost per task $5.4652; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI.

Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / max · reviewed 2026-09-29

Introducing GPT-6.1 Sollab self-report2026-09-29