RankingGPT-5.6 Terra

GPT-5.6 Terra

Data updated 12 Sept 2026

54 published benchmark measures · 17 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
GPT-5.6 Terra capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100GPT-5.6 Terra · Agentic: 40.9 · SupportedGPT-5.6 Terra · Hard reasoning: 52.1 · SupportedGPT-5.6 Terra · Coding: 61.0 · SupportedGPT-5.6 Terra · Multimodal: 48.9 · SupportedGPT-5.6 Terra · Long context: 50.9 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic5 families · 0 with independent evidence · Supported40.9

5 core families; 11 direct opponents across 3 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 34.8–49.2; 0/11 scenarios unsupported. Without one publisher: 39.2–45.3; 1/15 unsupported. Smoothing check: 38.7–44.0; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 38.6 observed win share
    Google, OpenAI · Source 1 Source 2
  • Browsecomp: 80.0 observed win share
    OpenAI · Source 1
  • AutomationBench: 50.0 observed win share
    OpenAI · Source 1
  • OSWorld: 50.0 observed win share
    OpenAI · Source 1
  • Toolathlon: 16.7 observed win share
    OpenAI · Source 1
Hard reasoning4 families · 2 with independent evidence · Supported52.1

4 core families; 58 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 49.6–57.1; 0/7 scenarios unsupported. Without one publisher: 36.0–53.0; 0/16 unsupported. Smoothing check: 51.1–52.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding3 families · 1 with independent evidence · Supported61.0

3 core families; 27 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 56.1–71.2; 0/7 scenarios unsupported. Without one publisher: 58.4–63.5; 0/20 unsupported. Smoothing check: 59.0–61.3; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Multimodal3 families · 0 with independent evidence · Supported48.9

      3 core families; 8 direct opponents across 3 labs.

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: 46.9–61.5; 0/8 scenarios unsupported. Without one publisher: 35.1–66.8; 1/8 unsupported. Smoothing check: 48.0–49.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Long context2 families · 0 with independent evidence · Supported50.9

      2 core families; 4 direct opponents across 2 labs.

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: 50.9–51.5; 2/5 scenarios unsupported. Without one publisher: 50.9–51.7; 2/5 unsupported. Smoothing check: 48.4–53.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

      Compare 7 effort levels across 68 benchmark/harness combinations →

      Reported effort · Not specified

      Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

      Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

      Inspect each result and its source ↓ · Download effort evidence

      Compare capability profiles →

      Score contributions and missing evidence

      17 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.

      Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

      Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

      Model information & shareable badge
      Lab
      OpenAI
      Catalog status
      active
      Availability
      Public provider catalog; account and region restrictions may apply
      Family
      GPT-5.6
      Released
      Context
      1,050,000 tokens
      License
      proprietary
      Model card
      https://developers.openai.com/api/docs/models/all
      Default Capability family coverage
      Documented-evidence family coverage/badge/gpt-5.6-terra.svg
      Benchmark scores & sources

      Original results, evaluation harnesses, and evidence behind this model.

      BenchmarkBucketScoreHarnessEvidenceSource-recorded date
      Agents' Last Exam · not specifiedSupporting evidence50.4%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ARC-AGI-2Hard reasoning83.9%ARC Prize verified

      Effort: Not specified

      contributes to capability
      official board2026-07-09
      ARC-AGI-3 · 3Hard reasoning0.8%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Artificial Analysis Coding Agent Index v1.1 · v1.1Supporting evidence77.4 index score
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Artificial Analysis Intelligence Index v4.1 · v4.1Supporting evidence55 index score
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      AutomationBench · not specifiedAgentic15.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      BenchCAD · not specifiedSupporting evidence62.3%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BenchCAD (python tool) · not specifiedSupporting evidence78.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Big Finance Bench · not specifiedSupporting evidence51%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BioMysteryBench · Human DifficultSupporting evidence49.4%
      Reported settings & source

      Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BioMysteryBench · Human SolvableSupporting evidence83.8%
      Reported settings & source

      Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      BrowseComp · not specifiedAgentic87.5%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Capture-the-Flag Challenges · not specifiedSupporting evidence91.8%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / Capture-the-Flag Challenges / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      CharXiv ReasoningMultimodal85.9%
      Reported settings & source

      No tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      DeepSWE · 1.1Coding69.6%
      Reported settings & source

      Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

      Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      DeepSWE v1.1Coding69.6%DeepSWE v1.1 reported

      Effort: Not specified

      contributes to capability
      official board2026-09-03
      DeepSWE v1.1 · v1.1Coding69.6%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      ExploitBench · not specifiedSupporting evidence52.9%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      ExploitGym · not specifiedSupporting evidence23.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; six-hour evaluation cap; alpha API latency rescaled to public API

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitGym / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence54.4%
      Reported settings & source

      Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

      First-party reported result; comparator results retain the source evaluation setup.

      Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      FrontierMath Tier 1-3 (v2) · v2Hard reasoning84.9%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      FrontierMath Tier 4 (v2) · v2Hard reasoning68.3%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GDP.PDFSupporting evidence29%
      Reported settings & source

      All-pass rate; allmodels selfcomputed byGoogle

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      gdp.pdf · not specifiedSupporting evidence24.7%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      GDPval-AA · 2Agentic1528
      Reported settings & source

      Artificial Analysis publicboard snapshot; effort as reported

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GDPval-AA v2 · v2Agentic1593
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GeneBench Pro · not specifiedSupporting evidence23.3%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      GPQA DiamondHard reasoning92.525%GPQA Diamond reported

      Effort: Not specified

      contributes to capability
      independent repro2026-09-11
      GPQA Diamond · not specifiedHard reasoning92.9%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GraphWalks BFS 1mil f1 · not specifiedLong context71.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GraphWalks BFS 256k f1 · not specifiedLong context76.9%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Harvey Legal Agent Benchmark · source release snapshot; version not specifiedSupporting evidence0.8%
      Reported settings & source

      Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

      First-party reported result; comparator results retain the source evaluation setup.

      Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench · not specifiedSupporting evidence57 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2285 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench · not specifiedSupporting evidence58.7 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2285 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-terra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Consensus · not specifiedSupporting evidence95.1 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2247 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Consensus · not specifiedSupporting evidence95.2 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2247 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-terra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Hard · not specifiedSupporting evidence32.7 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 2199 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Hard · not specifiedSupporting evidence34.3 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 2199 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-terra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional · not specifiedSupporting evidence57.7%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional · not specifiedSupporting evidence57.7 score (0-100)
      Reported settings & source

      length-adjusted; official HealthBench scoring

      Mean response length 3618 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HealthBench Professional · not specifiedSupporting evidence62.4 score (0-100)
      Reported settings & source

      unadjusted; official HealthBench scoring

      Mean response length 3618 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations.

      GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-terra / unadjusted · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      HLE-Verified · source release snapshot; version not specifiedHard reasoning51.1%
      Reported settings & source

      Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

      First-party reported result; comparator results retain the source evaluation setup.

      Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Internal Research Debugging Evaluation · not specifiedSupporting evidence67.8%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / Internal Research Debugging Evaluation / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      KernelGen 1P · not specifiedSupporting evidence49.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / KernelGen 1P / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LABBench · 2Supporting evidence81.2%
      Reported settings & source

      Selfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LifeSciBench · not specifiedSupporting evidence56%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      LVBench · staticMultimodal78.9%
      Reported settings & source

      No tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024

      Frame budgets differ; table labels Gemini3.8static explicitly.

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Management Consulting Tasks (Internal) · not specifiedSupporting evidence37.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      MedChemBench (Internal) · not specifiedSupporting evidence35%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      MMMU Pro (no tools) · not specifiedMultimodal80.7%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; no tools

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      MMMU Pro (with tools) · not specifiedMultimodal82%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; tools enabled

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (with tools) / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      NanoGPT · not specifiedSupporting evidence14.5%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / NanoGPT / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      OpenAI MRCR v2 8-needle 256K-512K · v2Long context89.6%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 256K-512K

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OpenAI MRCR v2 8-needle 512K-1M · v2Long context72.5%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; 8-needle 512K-1M

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OSWorld 2.0 · 2.0Agentic50.2%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      OSWorld partial · 2.0 task revision unverified gpt-5.6-terraAgentic50.2%
      Reported settings & source

      Partialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness

      Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports.

      Comparison limit: Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison.

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison.

      Reviewed 2026-09-06
      PostTrainBench Lite · not specifiedSupporting evidence51.5%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / PostTrainBench Lite / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      RSI Index · not specifiedSupporting evidence56.3%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / RSI Index / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SEC-Bench Pro · May2026 / public graderSupporting evidence57.7%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; public grader; May2026 JavaScript subset

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      lab self-report

      Supporting evidence outside the reviewed capability core

      Reviewed 2026-09-06
      SWE-Bench Pro · not specifiedCoding63.4%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench · 2.1Coding87.4%
      Reported settings & source

      Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench · 4.0Coding23.6%
      Reported settings & source

      Officialpublicboard highest scoring thinking level; nativeagents may differ

      Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Terminal-Bench 2.1 · 2.1Coding87.4%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      Toolathlon · not specifiedAgentic53.1%
      Reported settings & source

      Launch table reported configuration; per-cell reasoning effort unspecified

      Provider-published result; comparator measurements are not automatically independently reproduced.

      GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / GPT‑5.6 Terra · reviewed 2026-09-06

      Published configuration

      Effort: Not specified

      contributes to capability
      lab self-reportReviewed 2026-09-06
      GDPval-AAAgentic
      Humanity's Last ExamHard reasoning
      LiveCodeBenchCoding
      LMArena Text ArenaHuman pref
      MMLU-ProKnowledge
      OSWorld-VerifiedAgentic
      SWE-bench ProAgentic
      SWE-bench VerifiedAgentic
      Terminal-Bench 2.1Agentic