RankingClaude Opus 5

Claude Opus 5

Data updated 12 Sept 2026

65 published benchmark measures · 21 benchmark families contribute across 6 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Claude Opus 5 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Claude Opus 5 · Agentic: 85.9 · SupportedClaude Opus 5 · Hard reasoning: 79.7 · SupportedClaude Opus 5 · Coding: 92.5 · SupportedClaude Opus 5 · Human pref: 56.3 · SupportedClaude Opus 5 · Multimodal: 49.4 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic6 families · 2 with independent evidence · Supported85.9

6 core families; 52 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 82.4–92.9; 0/11 scenarios unsupported. Without one publisher: 80.7–91.9; 0/15 unsupported. Smoothing check: 77.0–90.1; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Hard reasoning5 families · 2 with independent evidence · Supported79.7

5 core families; 58 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 74.3–83.4; 0/7 scenarios unsupported. Without one publisher: 64.7–82.0; 0/16 unsupported. Smoothing check: 72.7–82.7; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding4 families · 1 with independent evidence · Supported92.5

4 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 83.2–93.7; 0/7 scenarios unsupported. Without one publisher: 84.7–93.9; 0/20 unsupported. Smoothing check: 85.5–95.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human pref1 families · 1 with independent evidence · Supported56.3

1 core families; 114 direct opponents across 19 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: 53.0–60.8; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Human preference (Arena): 96.5 observed win share
    lmarena.ai · Source 1
Knowledge2 families · 0 with independent evidence · PreliminaryUnknown

2 core families; 2 direct opponents across 1 labs. No connected comparison to the complete reference panel

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • Gmmlu: 0.0 observed win share
    Anthropic · Source 1
  • Milu: 0.0 observed win share
    Anthropic · Source 1
Multimodal3 families · 0 with independent evidence · Preliminary49.4

3 core families; 7 direct opponents across 3 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Long contextNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

    Compare 6 effort levels across 68 benchmark/harness combinations →

    Reported effort · Mixed settings

    Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

    Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

    Inspect each result and its source ↓ · Download effort evidence

    Compare capability profiles →

    Score contributions and missing evidence

    21 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.

    Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

    Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

    Model information & shareable badge
    Lab
    Anthropic
    Catalog status
    active
    Availability
    Public provider catalog; account and region restrictions may apply
    Family
    Claude Opus
    Released
    Context
    1,000,000 tokens
    License
    proprietary
    Model card
    https://platform.claude.com/docs/en/about-claude/model-deprecations
    Default Capability family coverage
    Documented-evidence family coverage/badge/claude-opus-5.svg
    Benchmark scores & sources

    Original results, evaluation harnesses, and evidence behind this model.

    BenchmarkBucketScoreHarnessEvidenceSource-recorded date
    AA-BriefcaseSupporting evidence1685
    Reported settings & source

    Artificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-BriefcaseSupporting evidence57.2%
    Reported settings & source

    Artificial Analysis; max effort; component rubric pass.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-BriefcaseSupporting evidence1980 rating
    Reported settings & source

    Artificial Analysis; max effort; component analytical quality.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    AA-BriefcaseSupporting evidence1572 rating
    Reported settings & source

    Artificial Analysis; max effort; component presentation.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Agentic IF Index · internalSupporting evidence59.1 index
    Reported settings & source

    Internal composite instruction-following evaluations; no fixed task count; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Agents' Last Exam · not specifiedSupporting evidence55.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ARC-AGI · 1Hard reasoning97.5%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI · 2Hard reasoning90.42%
    Reported settings & source

    ARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ARC-AGI-1 · 1Hard reasoning97.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-2Hard reasoning90.4%ARC Prize verified

    Effort: Not specified

    contributes to capability
    official board2026-07-24
    ARC-AGI-2 · 2Hard reasoning90.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ARC-AGI-3 · 3Hard reasoning30.2%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Artificial Analysis Coding Agent Index v1.4 · v1.4Supporting evidence68.1 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Artificial Analysis Intelligence IndexSupporting evidence63 AA Intelligence Index

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-08-12
    Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence63.1 index score
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    AutomationBenchAgentic26.9%
    Reported settings & source

    Private held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    AutomationBench · not specifiedAgentic26.9%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    AutomationBench · public v3Agentic50.3%
    Reported settings & source

    600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    BenchCAD · not specifiedSupporting evidence82.1%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled; three Anthropic evaluation modifications

    Provider-published result; comparator measurements are not automatically independently reproduced. Astra launch footnote5: Claude scores use three modifications described in the Fable5.1 system card; not same controlled setting.

    GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.366 voxel IoU
    Reported settings & source

    Random1000 of17900files; five runs; adaptive thinking max; no tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BenchCAD · Vision2Code 1000-file subsetSupporting evidence0.821 voxel IoU
    Reported settings & source

    Random1000 of17900files; five runs; adaptive thinking max; with tools; corrected camera prompt, raw shapes accepted, last code fence parsed.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BioMysteryBench · Human DifficultSupporting evidence51.8%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BioMysteryBench · Human DifficultSupporting evidence49.4%
    Reported settings & source

    Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BioMysteryBench · Human SolvableSupporting evidence91.4%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    BioMysteryBench · Human SolvableSupporting evidence90.1%
    Reported settings & source

    Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    BrowseCompAgentic90.8%BrowseComp reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-24
    BrowseComp · not specifiedAgentic90.8%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ChartographyMultimodal29.6%
    Reported settings & source

    100tasks; adaptive thinking max; five runs; no tools; tools condition has container, standard libraries and crop tool.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    ChartographyMultimodal83%
    Reported settings & source

    100tasks; adaptive thinking max; five runs; with tools; tools condition has container, standard libraries and crop tool.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    CharXiv ReasoningMultimodal83.7%
    Reported settings & source

    No tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    CursorBench · 3.2.0Supporting evidence70%
    Reported settings & source

    Cursor production agent harness; independently measured by Cursor; max effort.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    DeepSearchQAAgentic90.4 percent F1
    Reported settings & source

    900questions; common search backend/browser harness; answer-set F1; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    DeepSWE v1.1Coding73.6%DeepSWE v1.1 reported

    Effort: Not specified

    contributes to capability
    official board2026-09-03
    DeepSWE v1.1 · v1.1Coding73.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    ExploitBench · not specifiedSupporting evidence70%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    ExploitGym · not specifiedSupporting evidence22%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; launch-reported configuration

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence58.6%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode 1.1 Extended (score) · 1.1Supporting evidence63.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierCode 1.1 Main (score) · 1.1Supporting evidence53.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    FrontierMath Tier 4 (v2) · v2Hard reasoning73.2%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    FrontierSWE · 2Supporting evidence0.52 fraction
    Reported settings & source

    Proximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    GDP.PDFSupporting evidence37%
    Reported settings & source

    All-pass rate; allmodels selfcomputed byGoogle

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    GDPval-AAAgentic1735Artificial Analysis GDPval-AA

    Effort: Not specified

    contributes to capability
    official board2026-09-12
    GDPval-AA · 2Agentic1824
    Reported settings & source

    Artificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GDPval-AA · 2Agentic1824
    Reported settings & source

    Artificial Analysis publicboard snapshot; effort as reported

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GDPval-AA · 2Agentic1824
    Reported settings & source

    Artificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    GMMLUKnowledge92.5%
    Reported settings & source

    Mean accuracy42languages; adaptive max; one trial; no tools/custom system prompts.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    GPQA DiamondHard reasoning93.232%GPQA Diamond reported

    Effort: Not specified

    contributes to capability
    independent repro2026-09-11
    GPQA Diamond · not specifiedHard reasoning93.7%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Harvey Legal Agent Benchmark · source release snapshot; version not specifiedSupporting evidence6.7%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HealthBenchSupporting evidence67.1%
    Reported settings & source

    Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench ProfessionalSupporting evidence73.4%
    Reported settings & source

    Raw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench ProfessionalSupporting evidence59.8%
    Reported settings & source

    Length-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    HealthBench Professional (length-adjusted) · not specifiedSupporting evidence56.4%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted; OpenAI reproduction; GPT-5.4 grader; unclipped

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    HLE-Verified · source release snapshot; version not specifiedHard reasoning54.4%
    Reported settings & source

    Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

    First-party reported result; comparator results retain the source evaluation setup.

    Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Humanity's Last ExamHard reasoning56.6%HLE no tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-09-01
    Humanity's Last ExamHard reasoning56.3%HLE no tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-24
    Humanity's Last ExamHard reasoning63.6%HLE with tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-09-01
    Humanity's Last ExamHard reasoning64.7%HLE with tools

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-24
    Humanity’s Last ExamHard reasoning56.6%
    Reported settings & source

    Full2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

    Published configuration

    Effort: Auto thinking

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Humanity’s Last ExamHard reasoning63.6%
    Reported settings & source

    Full2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12

    Published configuration

    Effort: Auto thinking

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Humanity's Last Exam (w/ tools) · not specifiedHard reasoning63.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; tools enabled

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    JobBenchSupporting evidence65.7%
    Reported settings & source

    65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LABBench · 2Supporting evidence84.2%
    Reported settings & source

    Selfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    LMArena Text ArenaHuman pref1493LMArena Text

    Effort: Not specified

    contributes to capability
    official board2026-09-11
    LVBench · staticMultimodal75.4%
    Reported settings & source

    No tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 300

    Frame budgets differ; table labels Gemini3.8static explicitly.

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    MILUKnowledge92.1%
    Reported settings & source

    Mean accuracy11languages; adaptive max; five trials; no tools/custom system prompts.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OfficeQAAgentic78.1%
    Reported settings & source

    Extracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OfficeQA ProAgentic66.9%
    Reported settings & source

    Extracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Organic Chemistry · 2 revisedSupporting evidence65.7%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    OSWorld · 2.0 August2026 task releaseAgentic75.4%
    Reported settings & source

    partial pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

    Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld · 2.0 August2026 task releaseAgentic39.6%
    Reported settings & source

    strict pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero.

    Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offlineAgentic70.2%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld binary · 2.0 08.08Agentic31.4%
    Reported settings & source

    108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max

    Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld partial · 2.0 08.08Agentic68.3%
    Reported settings & source

    108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max

    Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    OSWorld partial · 2.0 August2026 fixed tasksAgentic75.4%
    Reported settings & source

    Partialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness

    Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports.

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Mixed source task revisions and best-of-three versus provider reporting prevent a uniform common-core unit.

    Reviewed 2026-09-06
    OSWorld-VerifiedAgentic83.39%OSWorld-Verified reported

    Effort: Not specified

    contributes to capability
    official board2026-08-01
    ProgramBench · 166 golden-task subsetSupporting evidence85.4%
    Reported settings & source

    mini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext.

    Hidden-test pass rate, not fraction of completely solved programs.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protein Design · Library RankingSupporting evidence48%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protein Design · Sequence Generation revised graderSupporting evidence42.4%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    ProteinGym · HardSupporting evidence47.7 percent rank correlation
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protocols · TroubleshootingSupporting evidence61.1%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    Protocols · Understanding network-restrictedSupporting evidence80%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SingleCellBenchSupporting evidence60.6%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SpatialBench · VerifiedSupporting evidence72.5%
    Reported settings & source

    Anthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction.

    Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SRE-Bench · 262 binaries / 19 programsSupporting evidence12.5%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-Atlas Codebase QnASupporting evidence52.7%
    Reported settings & source

    124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-06
    SWE-bench MultilingualCoding89.5%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-bench MultimodalSupporting evidence59.4%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    lab self-report

    Supporting evidence outside the reviewed capability core

    Reviewed 2026-09-12
    SWE-bench ProCoding79.2%SWE-bench Pro reported

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-24
    SWE-bench ProCoding79.2%
    Reported settings & source

    Anthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    SWE-bench VerifiedCoding96%SWE-bench Verified official

    Effort: Not specified

    lab self-report

    Legacy self-report lacks a reviewed comparison configuration

    2026-07-24
    Terminal-Bench · 2.1Coding89.1%
    Reported settings & source

    Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench · 2.1Coding86.7%
    Reported settings & source

    89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Anthropic (exact harness revision not specified)

    Native harnesses differ; not a Terminus2-only comparison.

    Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    lab self-report

    Different native coding harnesses are mixed; exact shared agent configuration is not established.

    Reviewed 2026-09-06
    Terminal-Bench · 4.0Coding52.3%
    Reported settings & source

    Claude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns.

    Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12

    Published configuration

    Effort: Max

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Terminal-Bench · 4.0Coding51.8%
    Reported settings & source

    Officialpublicboard highest scoring thinking level; nativeagents may differ

    Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench 4.0 · 4.0Coding52.6%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench Science 0.1 · not specifiedHard reasoning30%
    Reported settings & source

    Maximum reported across reasoning efforts; OpenAI research environment or API

    Provider-published result; comparator measurements are not automatically independently reproduced.

    GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / Claude Opus 5 · reviewed 2026-09-06

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-06
    Terminal-Bench-Science · 0.1Coding29%
    Reported settings & source

    70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12

    Published configuration

    Effort: Not specified

    contributes to capability
    lab self-reportReviewed 2026-09-12
    Toolathlon · Verified June2026Agentic80.6%
    Reported settings & source

    Pass@1;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-12
    Toolathlon · Verified June2026Agentic87%
    Reported settings & source

    Pass@3;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-12
    Toolathlon · Verified June2026Agentic73.1%
    Reported settings & source

    Pass³;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    No matched opponent in this evaluation unit

    Reviewed 2026-09-12
    Toolathlon · Verified June2026Agentic23.5 turns
    Reported settings & source

    average turns;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data.

    Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12

    Published configuration

    Effort: Max

    lab self-report

    Metric turns is outside the task-success core

    Reviewed 2026-09-12
    VulcanBench v3Supporting evidence78.3%VulcanBench v3 bare-bones API · Report 10 · high

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-26
    VulcanBench v3Supporting evidence87%VulcanBench v3 bare-bones API · Report 10 · low

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-26
    VulcanBench v3Supporting evidence82.6%VulcanBench v3 bare-bones API · Report 10 · medium

    Effort: Not specified

    official board

    Supporting evidence outside the reviewed capability core

    2026-07-26
    LiveCodeBenchCoding
    MMLU-ProKnowledge
    Terminal-Bench 2.1Agentic