RankingGPT-6 Astra
GPT-6 Astra
49 published benchmark measures · 12 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 6 effort levels across 65 benchmark/harness combinations →
Reported effort · Best across efforts + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Best across efforts: 34 observations
- Not specified: 26 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
12 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- OpenAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- GPT-6
- Released
- —
- Context
- 1,050,000 tokens
- License
- proprietary
- Model card
- https://developers.openai.com/api/docs/models/all
- Default Capability family coverage
/badge/gpt-6-astra.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agents' Last Exam · not specified | Supporting evidence | 59.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-1 · 1 | Hard reasoning | 98.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 95% | ARC Prize verifiedcontributes to capability | official board | 2026-09-02 |
| ARC-AGI-2 · 2 | Hard reasoning | 95%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-3 · 3 | Hard reasoning | 99.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; Responses API harness with two documented setting changes Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.4 · v1.4 | Supporting evidence | 67 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1.1 · v4.1.1 | Supporting evidence | 61.2 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · not specified | Agentic | 41.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 95.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · not specified | Agentic | 91.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Coronavirus-ACE2 Cell-Entry Screen · not specified | Supporting evidence | 0.42 composite scoreReported settings & sourceSystem-card reported configuration; reasoning effort unspecified Observed named-model performance; helpful-only checkpoint omitted as distinct noncatalog identity. GPT-6 Astra System Card · Section10.1.1.2.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 74.1% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 74.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · not specified | Supporting evidence | 100%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitBench (June-Aug 2026) · June-Aug2026 | Supporting evidence | 39%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; 20 vulnerabilities / 13 Chrome releases Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench (June-Aug 2026) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 42.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; v1 offline environment; no runtime package installation; token capped, no wall-clock cap Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Extended (score) · 1.1 | Supporting evidence | 64.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; Codex-style developer instruction on tests, reuse and repository conventions Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Main (score) · 1.1 | Supporting evidence | 53.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; Codex-style developer instruction on tests, reuse and repository conventions Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 97.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · v13 | Supporting evidence | 37.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / GeneBench Pro / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 96.061% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 96%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · not specified | Hard reasoning | 94.9%Reported settings & sourceLower-cost setting; exact effort not specified Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra: A new generation of intelligence · GPQA Diamond chart caption · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 58.1 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2258 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-6-astra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 59.7 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2258 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-6-astra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Consensus · not specified | Supporting evidence | 95.8 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2237 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-6-astra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Consensus · not specified | Supporting evidence | 95.9 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2237 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-6-astra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 36.3 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2192 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-6-astra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 37.8 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2192 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-6-astra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 63.4 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 4097 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-6-astra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 69.5 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 4097 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-6-astra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional (length-adjusted) · not specified | Supporting evidence | 63.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Humanity's Last Exam (w/ tools) · not specified | Hard reasoning | 57.2%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Internal Data Science Tasks · not specified | Supporting evidence | 40.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Internal Data Science Tasks / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Database Migration Tasks · not specified | Supporting evidence | 63.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Design Tasks · not specified | Supporting evidence | 50%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Internal Design Tasks / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Internal Research Debugging Evaluation · 41 research bugs; 6 alignment-auditing tasks | Supporting evidence | 78.05%Reported settings & sourceSystem-card reported configuration; reasoning effort unspecified Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.3.1 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LifeSciBench · Gold v1 | Supporting evidence | 60.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / LifeSciBench / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MedChemBench (Internal) · not specified | Supporting evidence | 49.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / MedChemBench (Internal) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| No-CoT math time horizon · not specified | Supporting evidence | 30.9 minutesReported settings & sourceUK AISI; single forward pass; no chain of thought Time-horizon estimate; possible contamination noted by evaluator. Higher means harder tasks solved, not slower inference. GPT-6 Astra System Card · Section9.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 256K-512K · v2 | Long context | 100%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 256K-512K Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 512K-1M · v2 | Long context | 96.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 512K-1M Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenScore String Quartets (1 - OMR-NED) · not specified | Supporting evidence | 0.84 1 - OMR-NEDReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / OpenScore String Quartets (1 - OMR-NED) / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 72.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Phage-plasmid Co-evolution · not specified | Supporting evidence | 13.13 negative log-likelihoodReported settings & sourceSystem-card reported configuration; reasoning effort unspecified Lower is better. Helpful-only checkpoint omitted as distinct noncatalog identity. GPT-6 Astra System Card · Section10.1.1.2.4 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ProtocolQA Open-Ended · 108 questions | Supporting evidence | 41.36%Reported settings & sourceobserved performance Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / ProtocolQA Open-Ended · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ProtocolQA Open-Ended · 108 questions | Supporting evidence | 45.37%Reported settings & sourcerefusal-adjusted upper estimate; refusals counted as successes Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / ProtocolQA Open-Ended · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Sandbox Bench · September2026 internal | Supporting evidence | 45.5%Reported settings & source22 isolated CTF-style targets; protected-flag success metric Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.2.4 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ScreenSpot-Pro (no tools) · not specified | Multimodal | 92.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; no tools Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / ScreenSpot-Pro (no tools) / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SEC-Bench Pro · May2026 / revised root-cause grader | Supporting evidence | 85.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; May2026 JavaScript subset,183 vulnerabilities; revised agent root-cause grader Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SEC-Bench Pro / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SHP2 Protein Function Prediction · 3 unpublished assay datasets | Supporting evidence | 0.4 mean R-squaredReported settings & sourceSystem-card reported configuration; reasoning effort unspecified Production-named model result; separate helpful-only checkpoint omitted because no exact catalog identity. GPT-6 Astra System Card · Section10.1.1.2.2 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SRE-Bench · 262 binaries / 19 programs | Supporting evidence | 88%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / GPT‑6 Astra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SRE-Bench · 262 binaries / 19 programs | Supporting evidence | 99.2%Reported settings & sourcepass@4; four independent trials; all six objectives required; reduced production safeguards Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.2.3 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Tacit Knowledge and Troubleshooting · 60 questions | Supporting evidence | 63.33%Reported settings & sourceobserved performance Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / Tacit Knowledge and Troubleshooting · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Tacit Knowledge and Troubleshooting · 60 questions | Supporting evidence | 90%Reported settings & sourcerefusal-adjusted upper estimate; refusals counted as successes Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / Tacit Knowledge and Troubleshooting · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 4.0 · 4.0 | Coding | 57.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Science 0.1 · 0.1 | Hard reasoning | 61.1%Reported settings & sourceLower-cost setting; exact effort not specified Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra: A new generation of intelligence · Terminal-Bench Science chart caption · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Science 0.1 · not specified | Hard reasoning | 64.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / GPT‑6 Astra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| TroubleshootingBench · 156 questions / 52 protocols | Supporting evidence | 48.44%Reported settings & sourceobserved performance Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / TroubleshootingBench · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| TroubleshootingBench · 156 questions / 52 protocols | Supporting evidence | 63.46%Reported settings & sourcerefusal-adjusted upper estimate; refusals counted as successes Provider-published result; preserves source setting without claiming a controlled cross-provider comparison. GPT-6 Astra System Card · Section10.1.1.1 / TroubleshootingBench · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |