RankingGPT-5.6 Terra
GPT-5.6 Terra
54 published benchmark measures · 17 benchmark families contribute across 5 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 7 effort levels across 68 benchmark/harness combinations →
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 64 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
17 contributing families across 5 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- OpenAI
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- GPT-5.6
- Released
- —
- Context
- 1,050,000 tokens
- License
- proprietary
- Model card
- https://developers.openai.com/api/docs/models/all
- Default Capability family coverage
/badge/gpt-5.6-terra.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agents' Last Exam · not specified | Supporting evidence | 50.4%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Agents' Last Exam / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 83.9% | ARC Prize verifiedcontributes to capability | official board | 2026-07-09 |
| ARC-AGI-3 · 3 | Hard reasoning | 0.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Abstract Reasoning table / ARC-AGI-3 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.1 · v1.1 | Supporting evidence | 77.4 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Artificial Analysis Coding Agent Index v1.1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1 · v4.1 | Supporting evidence | 55 index scoreReported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Artificial Analysis Intelligence Index v4.1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · not specified | Agentic | 15.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / AutomationBench / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 62.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD (python tool) · not specified | Supporting evidence | 78.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; Python tool enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BenchCAD (python tool) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Big Finance Bench · not specified | Supporting evidence | 51%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Big Finance Bench / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Difficult | Supporting evidence | 49.4%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Solvable | Supporting evidence | 83.8%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · not specified | Agentic | 87.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / BrowseComp / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Capture-the-Flag Challenges · not specified | Supporting evidence | 91.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / Capture-the-Flag Challenges / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CharXiv Reasoning | Multimodal | 85.9%Reported settings & sourceNo tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · 1.1 | Coding | 69.6%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 69.6% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 69.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / DeepSWE v1.1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · not specified | Supporting evidence | 52.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; ExploitBench API harness; five seeds; reasoning continuity Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitBench / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 23.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; six-hour evaluation cap; alpha API latency rescaled to public API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / ExploitGym / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 54.4%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 1-3 (v2) · v2 | Hard reasoning | 84.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 1-3 (v2) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 68.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / FrontierMath Tier 4 (v2) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDP.PDF | Supporting evidence | 29%Reported settings & sourceAll-pass rate; allmodels selfcomputed byGoogle Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| gdp.pdf · not specified | Supporting evidence | 24.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / gdp.pdf / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1528Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA v2 · v2 | Agentic | 1593Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / GDPval-AA v2 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GeneBench Pro · not specified | Supporting evidence | 23.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / GeneBench Pro / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 92.525% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 92.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Academic table / GPQA Diamond / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GraphWalks BFS 1mil f1 · not specified | Long context | 71.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 1mil f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 1mil f1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GraphWalks BFS 256k f1 · not specified | Long context | 76.9%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; BFS 256k f1 Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / GraphWalks BFS 256k f1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey Legal Agent Benchmark · source release snapshot; version not specified | Supporting evidence | 0.8%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 57 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2285 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench · not specified | Supporting evidence | 58.7 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2285 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench / gpt-5.6-terra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Consensus · not specified | Supporting evidence | 95.1 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2247 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Consensus · not specified | Supporting evidence | 95.2 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2247 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Consensus / gpt-5.6-terra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 32.7 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 2199 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Hard · not specified | Supporting evidence | 34.3 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 2199 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Hard / gpt-5.6-terra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 57.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / HealthBench Professional / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 57.7 score (0-100)Reported settings & sourcelength-adjusted; official HealthBench scoring Mean response length 3618 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-terra / length-adjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · not specified | Supporting evidence | 62.4 score (0-100)Reported settings & sourceunadjusted; official HealthBench scoring Mean response length 3618 characters. Adjustment centered on2000characters; adjusted and unadjusted values remain distinct configurations. GPT-6 Astra System Card · Table6 / HealthBench Professional / gpt-5.6-terra / unadjusted · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE-Verified · source release snapshot; version not specified | Hard reasoning | 51.1%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Internal Research Debugging Evaluation · not specified | Supporting evidence | 67.8%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / Internal Research Debugging Evaluation / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| KernelGen 1P · not specified | Supporting evidence | 49.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / KernelGen 1P / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LABBench · 2 | Supporting evidence | 81.2%Reported settings & sourceSelfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LifeSciBench · not specified | Supporting evidence | 56%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / LifeSciBench / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LVBench · static | Multimodal | 78.9%Reported settings & sourceNo tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024 Frame budgets differ; table labels Gemini3.8static explicitly. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Management Consulting Tasks (Internal) · not specified | Supporting evidence | 37.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Professional table / Management Consulting Tasks (Internal) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MedChemBench (Internal) · not specified | Supporting evidence | 35%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Science And Health table / MedChemBench (Internal) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MMMU Pro (no tools) · not specified | Multimodal | 80.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; no tools Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (no tools) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU Pro (with tools) · not specified | Multimodal | 82%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Multimodal table / MMMU Pro (with tools) / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| NanoGPT · not specified | Supporting evidence | 14.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / NanoGPT / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 256K-512K · v2 | Long context | 89.6%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; 8-needle 256K-512K Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 256K-512K / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OpenAI MRCR v2 8-needle 512K-1M · v2 | Long context | 72.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; 8-needle 512K-1M Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Long Context table / OpenAI MRCR v2 8-needle 512K-1M / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld 2.0 · 2.0 | Agentic | 50.2%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Computer Use table / OSWorld 2.0 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 task revision unverified gpt-5.6-terra | Agentic | 50.2%Reported settings & sourcePartialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports. Comparison limit: Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| PostTrainBench Lite · not specified | Supporting evidence | 51.5%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / PostTrainBench Lite / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| RSI Index · not specified | Supporting evidence | 56.3%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Self-Improvement table / RSI Index / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SEC-Bench Pro · May2026 / public grader | Supporting evidence | 57.7%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified; reduced or absent production safeguards; public grader; May2026 JavaScript subset Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Cybersecurity table / SEC-Bench Pro / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Bench Pro · not specified | Coding | 63.4%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / SWE-Bench Pro / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 87.4%Reported settings & sourceTerminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 4.0 | Coding | 23.6%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.1 · 2.1 | Coding | 87.4%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Coding table / Terminal-Bench 2.1 / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Toolathlon · not specified | Agentic | 53.1%Reported settings & sourceLaunch table reported configuration; per-cell reasoning effort unspecified Provider-published result; comparator measurements are not automatically independently reproduced. GPT-5.6: Frontier intelligence that scales with your ambition · Tool Use table / Toolathlon / GPT‑5.6 Terra · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |