RankingTerminal-Bench Science 0.1 · 0.1
Terminal-Bench Science 0.1 · 0.1
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 0.1
- Display harness
- GPT-6 Astra: A new generation of intelligence
- Board
- https://openai.com/index/gpt-6-astra/
Compare published benchmark results with category weights →
Models
1–17 of 17 entries
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-6 AstraOpenAI | 61.1%Reported settings & sourceLower-cost setting; exact effort not specified Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra: A new generation of intelligence · Terminal-Bench Science chart caption · reviewed 2026-09-06 | GPT-6 Astra: A new generation of intelligence | lab self-report | 2026-09-06 |
| Claude Opus 5.5Anthropic | 63.33%Reported settings & sourceReported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. Reported combined setup: Opus 5.5 w/ fallbacks. Fallback behavior must not be attributed to the base model alone. Provider-published lab self-report. Published chart cost per task $23.2108; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Comparison limit: Combined fallback setup is not the named base model at one effort setting. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / Opus 5.5 w/ fallbacks / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 55.43%Reported settings & sourceReported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $11.4061; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 57.43%Reported settings & sourceReported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $12.3383; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 62%Reported settings & sourceReported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $14.9534; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 60.86%Reported settings & sourceReported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $15.7554; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 AstraOpenAI | 68.1%Reported settings & sourceReported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $23.7974; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Astra / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 9.17%Reported settings & sourceReported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.9986; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 14.49%Reported settings & sourceReported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $4.4068; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 14.61%Reported settings & sourceReported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $4.6325; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 25.29%Reported settings & sourceReported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $6.7698; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6 SolOpenAI | 27.59%Reported settings & sourceReported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $12.1803; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 43.71%Reported settings & sourceReported reasoning effort low. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $1.7918; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / low · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 47.56%Reported settings & sourceReported reasoning effort medium. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.3386; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / medium · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 51.14%Reported settings & sourceReported reasoning effort high. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.7594; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / high · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 53.71%Reported settings & sourceReported reasoning effort xhigh. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $2.8922; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / xhigh · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |
| GPT-6.1 SolOpenAI | 57.02%Reported settings & sourceReported reasoning effort max. Scientific research workflows using code and terminal tools, including analysis, simulation and model fitting. OpenAI research environment or API; same named metric and evaluation chart. No fallback is reported for this OpenAI configuration. Provider-published lab self-report. Published chart cost per task $5.4652; source-specific cost does not establish consensus workload efficiency. Competitor results are republished from public reports; the page does not establish independent reproduction by OpenAI. Introducing GPT-6.1 Sol · Chart: Terminal-Bench Science 0.1 / GPT-6.1 Sol / max · reviewed 2026-09-29 | Introducing GPT-6.1 Sol | lab self-report | 2026-09-29 |