RankingGPT-6 Astra

GPT-6 Astra

Data updated 12 Sept 2026

49 published benchmark measures · 12 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-6 Astra capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-6 Astra · Agentic: 97.7 · SupportedGPT-6 Astra · Hard reasoning: 94.6 · SupportedGPT-6 Astra · Coding: 97.7 · SupportedGPT-6 Astra · Multimodal: 94.6 · PreliminaryGPT-6 Astra · Long context: 94.4 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic3 families · 0 with independent evidence · Supported97.7

3 core families; 4 direct opponents across 2 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 97.5–98.7; 2/11 scenarios unsupported. Without one publisher: 97.0–98.6; 1/15 unsupported. Smoothing check: 91.2–99.4; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Browsecomp: 100.0 observed win share
    OpenAI · Source 1
  • AutomationBench: 100.0 observed win share
    OpenAI · Source 1
  • OSWorld: 100.0 observed win share
    OpenAI · Source 1
Hard reasoning5 families · 2 with independent evidence · Supported94.6

5 core families; 57 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 89.8–96.8; 0/7 scenarios unsupported. Without one publisher: 83.5–95.0; 0/16 unsupported. Smoothing check: 88.4–96.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 1 with independent evidence · Supported97.7

2 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 97.4–97.7; 2/7 scenarios unsupported. Without one publisher: 95.7–98.1; 1/20 unsupported. Smoothing check: 91.6–99.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • DeepSWE: 100.0 observed win share
    OpenAI, deepswe.datacurve.ai · Source 1 Source 2
  • Terminal-Bench: 100.0 observed win share
    OpenAI · Source 1
Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Multimodal1 families · 0 with independent evidence · Preliminary94.6

      1 core families; 1 direct opponents across 1 labs. Needs broader benchmark and opponent support

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      • ScreenSpot: 100.0 observed win share
        OpenAI · Source 1
      Long context1 families · 0 with independent evidence · Preliminary94.4

      1 core families; 1 direct opponents across 1 labs. Needs broader benchmark and opponent support

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 6 effort levels across 65 benchmark/harness combinations →

      Reported effort · Best across efforts + unspecified

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      12 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      OpenAI
      Catalog status
      active
      Availability
      Public provider catalog; account and region restrictions may apply
      Family
      GPT-6
      Released
      Context
      1,050,000 tokens
      License
      proprietary
      Model card
      https://developers.openai.com/api/docs/models/all
      Default Capability family coverage
      Documented-evidence family coverage/badge/gpt-6-astra.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      Agents' Last Exam · not specifiedSupporting evidence59.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ARC-AGI-1 · 1Hard reasoning98.5%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      ARC-AGI-2Hard reasoning95%ARC Prize verified

      Effort: Not specified

      contributes to capability
      official board2026-09-02
      ARC-AGI-2 · 2Hard reasoning95%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      ARC-AGI-3 · 3Hard reasoning99.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; Responses API harness with two documented setting changes

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-06
      Artificial Analysis Coding Agent Index v1.4 · v1.4Supporting evidence67 index score
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence61.2 index score
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      AutomationBench · not specifiedAgentic41.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      BenchCAD · not specifiedSupporting evidence95.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BrowseComp · not specifiedAgentic91.5%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Coronavirus-ACE2 Cell-Entry Screen · not specifiedSupporting evidence0.42 composite score
      Reported settings & source

      System-card reported configuration; reasoning effort unspecified

      Observed named-model performance; helpful-only checkpoint omitted as distinct noncatalog identity.

      GPT-6 Astra System Card · Section10.1.1.2.3 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      DeepSWE v1.1Coding74.1%DeepSWE v1.1 reported

      Effort: Not specified

      contributes to capability
      official board2026-09-03
      DeepSWE v1.1 · v1.1Coding74.1%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      ExploitBench · not specifiedSupporting evidence100%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ExploitBench (June-Aug 2026) · June-Aug2026Supporting evidence39%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; 20 vulnerabilities / 13 Chrome releases

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench (June-Aug 2026) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ExploitGym · not specifiedSupporting evidence42.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; v1 offline environment; no runtime package installation; token capped, no wall-clock cap

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierCode 1.1 Extended (score) · 1.1Supporting evidence64.5%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; Codex-style developer instruction on tests, reuse and repository conventions

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierCode 1.1 Main (score) · 1.1Supporting evidence53.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; Codex-style developer instruction on tests, reuse and repository conventions

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierMath Tier 4 (v2) · v2Hard reasoning97.6%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GeneBench Pro · v13Supporting evidence37.1%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Science And Health table / GeneBench Pro / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      GPQA DiamondHard reasoning96.061%GPQA Diamond reported

      Effort: Not specified

      contributes to capability
      independent repro2026-09-11
      GPQA Diamond · not specifiedHard reasoning96%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GPQA Diamond · not specifiedHard reasoning94.9%
      Reported settings & source

      Lower-cost setting; exact effort not specified

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra: A new generation of intelligence · GPQA Diamond chart caption · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-06
      HealthBench · not specifiedSupporting evidence58.1 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2258 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench / gpt-6-astra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench · not specifiedSupporting evidence59.7 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2258 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench / gpt-6-astra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Consensus · not specifiedSupporting evidence95.8 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2237 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-6-astra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Consensus · not specifiedSupporting evidence95.9 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2237 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-6-astra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Hard · not specifiedSupporting evidence36.3 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2192 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-6-astra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Hard · not specifiedSupporting evidence37.8 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2192 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-6-astra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional · not specifiedSupporting evidence63.4 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 4097 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-6-astra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional · not specifiedSupporting evidence69.5 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 4097 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-6-astra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional (length-adjusted) · not specifiedSupporting evidence63.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Humanity's Last Exam (w/ tools) · not specifiedHard reasoning57.2%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Internal Data Science Tasks · not specifiedSupporting evidence40.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / Internal Data Science Tasks / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Internal Database Migration Tasks · not specifiedSupporting evidence63.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / Internal Database Migration Tasks / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Internal Design Tasks · not specifiedSupporting evidence50%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / Internal Design Tasks / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Internal Research Debugging Evaluation · 41 research bugs; 6 alignment-auditing tasksSupporting evidence78.05%
      Reported settings & source

      System-card reported configuration; reasoning effort unspecified

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.3.1 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LifeSciBench · Gold v1Supporting evidence60.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Science And Health table / LifeSciBench / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      MedChemBench (Internal) · not specifiedSupporting evidence49.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Science And Health table / MedChemBench (Internal) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      No-CoT math time horizon · not specifiedSupporting evidence30.9 minutes
      Reported settings & source

      UK AISI; single forward pass; no chain of thought

      Time-horizon estimate; possible contamination noted by evaluator. Higher means harder tasks solved, not slower inference.

      GPT-6 Astra System Card · Section9.3 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      OpenAI MRCR v2 8-needle 256K-512K · v2Long context100%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 256K-512K

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OpenAI MRCR v2 8-needle 512K-1M · v2Long context96.3%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; 8-needle 512K-1M

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OpenScore String Quartets (1 - OMR-NED) · not specifiedSupporting evidence0.84 1 - OMR-NED
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Professional table / OpenScore String Quartets (1 - OMR-NED) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic72.6%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Phage-plasmid Co-evolution · not specifiedSupporting evidence13.13 negative log-likelihood
      Reported settings & source

      System-card reported configuration; reasoning effort unspecified

      Lower is better. Helpful-only checkpoint omitted as distinct noncatalog identity.

      GPT-6 Astra System Card · Section10.1.1.2.4 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ProtocolQA Open-Ended · 108 questionsSupporting evidence41.36%
      Reported settings & source

      observed performance

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / ProtocolQA Open-Ended · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ProtocolQA Open-Ended · 108 questionsSupporting evidence45.37%
      Reported settings & source

      refusal-adjusted upper estimate; refusals counted as successes

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / ProtocolQA Open-Ended · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Sandbox Bench · September2026 internalSupporting evidence45.5%
      Reported settings & source

      22 isolated CTF-style targets; protected-flag success metric

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.2.4 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ScreenSpot-Pro (no tools) · not specifiedMultimodal92.7%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; no tools

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Computer Use table / ScreenSpot-Pro (no tools) / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      SEC-Bench Pro · May2026 / revised root-cause graderSupporting evidence85.4%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; May2026 JavaScript subset,183 vulnerabilities; revised agent root-cause grader

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SEC-Bench Pro / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SHP2 Protein Function Prediction · 3 unpublished assay datasetsSupporting evidence0.4 mean R-squared
      Reported settings & source

      System-card reported configuration; reasoning effort unspecified

      Production-named model result; separate helpful-only checkpoint omitted because no exact catalog identity.

      GPT-6 Astra System Card · Section10.1.1.2.2 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SRE-Bench · 262 binaries / 19 programsSupporting evidence88%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SRE-Bench · 262 binaries / 19 programsSupporting evidence99.2%
      Reported settings & source

      pass@4; four independent trials; all six objectives required; reduced production safeguards

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.2.3 · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Tacit Knowledge and Troubleshooting · 60 questionsSupporting evidence63.33%
      Reported settings & source

      observed performance

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / Tacit Knowledge and Troubleshooting · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Tacit Knowledge and Troubleshooting · 60 questionsSupporting evidence90%
      Reported settings & source

      refusal-adjusted upper estimate; refusals counted as successes

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / Tacit Knowledge and Troubleshooting · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Terminal-Bench 4.0 · 4.0Coding57.9%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench Science 0.1 · 0.1Hard reasoning61.1%
      Reported settings & source

      Lower-cost setting; exact effort not specified

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra: A new generation of intelligence · Terminal-Bench Science chart caption · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      No matched opponent in this evaluation unit

      Reviewed 2026-09-06
      Terminal-Bench Science 0.1 · not specifiedHard reasoning64.6%
      Reported settings & source

      Maximum reported across reasoning efforts; OpenAI research environment or API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / GPT‑6 Astra · reviewed 2026-09-06

      Published configuration

      Effort: Best across efforts

      contributes to capability
      lab self-reportReviewed 2026-09-06
      TroubleshootingBench · 156 questions / 52 protocolsSupporting evidence48.44%
      Reported settings & source

      observed performance

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / TroubleshootingBench · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      TroubleshootingBench · 156 questions / 52 protocolsSupporting evidence63.46%
      Reported settings & source

      refusal-adjusted upper estimate; refusals counted as successes

      Provider-published result; preserves source setting without claiming a controlled cross-provider comparison.

      GPT-6 Astra System Card · Section10.1.1.1 / TroubleshootingBench · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      GDPval-AAAgentic
      Humanity's Last ExamHard reasoning
      LiveCodeBenchCoding
      LMArena Text ArenaHuman pref
      MMLU-ProKnowledge
      OSWorld-VerifiedAgentic
      SWE-bench ProAgentic
      SWE-bench VerifiedAgentic
      Terminal-Bench 2.1Agentic