RankingGemini 3.8 Flash

Gemini 3.8 Flash

Data updated 12 Sept 2026

22 published benchmark measures · 7 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Gemini 3.8 Flash capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Gemini 3.8 Flash · Agentic: 45.6 · PreliminaryGemini 3.8 Flash · Hard reasoning: 89.2 · SupportedGemini 3.8 Flash · Coding: 85.2 · SupportedGemini 3.8 Flash · Multimodal: 93.3 · Supported

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic1 families · 0 with independent evidence · Preliminary45.6

1 core families; 5 direct opponents across 3 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 40.0 observed win share
    Google · Source 1
Hard reasoning2 families · 2 with independent evidence · Supported89.2

2 core families; 29 direct opponents across 12 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 87.3–89.6; 2/7 scenarios unsupported. Without one publisher: 78.7–93.3; 0/16 unsupported. Smoothing check: 79.8–93.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Coding2 families · 1 with independent evidence · Supported85.2

2 core families; 28 direct opponents across 9 labs.

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: 83.8–87.1; 2/7 scenarios unsupported. Without one publisher: 80.0–87.2; 0/20 unsupported. Smoothing check: 77.3–89.2; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

Human prefNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    KnowledgeNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Multimodal2 families · 0 with independent evidence · Supported93.3

      2 core families; 5 direct opponents across 3 labs.

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: 92.7–94.5; 2/8 scenarios unsupported. Without one publisher: 76.9–95.4; 1/8 unsupported. Smoothing check: 84.8–96.9; 0/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      • CharXiv: 100.0 observed win share
        Google · Source 1
      • Lvbench: 100.0 observed win share
        Google · Source 1
      Long contextNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

        Compare 3 effort levels across 58 benchmark/harness combinations →

        Reported effort · High thinking + unspecified

        Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

        Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

        Inspect each result and its source ↓ · Download effort evidence

        Compare capability profiles →

        Score contributions and missing evidence

        7 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.

        Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

        Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

        Model information & shareable badge
        Lab
        Google
        Catalog status
        active
        Availability
        Public provider catalog; account and region restrictions may apply
        Family
        Gemini 3
        Released
        Context
        License
        proprietary
        Model card
        https://ai.google.dev/gemini-api/docs/models
        Default Capability family coverage
        Documented-evidence family coverage/badge/gemini-3.8-flash.svg
        Benchmark scores & sources

        Original results, evaluation harnesses, and evidence behind this model.

        BenchmarkBucketScoreHarnessEvidenceSource-recorded date
        Artificial Analysis Coding Agent Index v1.4 · v1.4Supporting evidence61.2 index score
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        Artificial Analysis Intelligence Index v4.1.1 · v4.1.1Supporting evidence58.7 index score
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        BioMysteryBench · Human DifficultSupporting evidence56.5%
        Reported settings & source

        Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        BioMysteryBench · Human SolvableSupporting evidence88.8%
        Reported settings & source

        Linuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        CharXiv ReasoningMultimodal86.2%
        Reported settings & source

        No tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        DeepSWE · 1.1Coding73.7%
        Reported settings & source

        Datacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking

        Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted.

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: High thinking

        contributes to capability
        lab self-reportReviewed 2026-09-06
        DeepSWE v1.1Coding73.8%DeepSWE v1.1 reported

        Effort: Not specified

        contributes to capability
        official board2026-09-03
        DeepSWE v1.1 · v1.1Coding73.8%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        Finance Agent v2 · source release snapshot; version not specifiedSupporting evidence61.4%
        Reported settings & source

        Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

        First-party reported result; comparator results retain the source evaluation setup.

        Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        FrontierCode 1.1 Extended (score) · 1.1Supporting evidence56.3%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        FrontierCode 1.1 Main (score) · 1.1Supporting evidence43.6%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        GDP.PDFSupporting evidence35%
        Reported settings & source

        All-pass rate; allmodels selfcomputed byGoogle

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        GDPval-AA · 2Agentic1545
        Reported settings & source

        Artificial Analysis publicboard snapshot; effort as reported

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        GPQA DiamondHard reasoning95.253%GPQA Diamond reported

        Effort: Not specified

        contributes to capability
        independent repro2026-09-11
        GPQA Diamond · not specifiedHard reasoning95.3%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        Harvey Legal Agent Benchmark · source release snapshot; version not specifiedSupporting evidence10%
        Reported settings & source

        Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

        First-party reported result; comparator results retain the source evaluation setup.

        Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        HealthBench Professional (length-adjusted) · not specifiedSupporting evidence52.1%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        HLE-Verified · source release snapshot; version not specifiedHard reasoning54.9%
        Reported settings & source

        Google launch chart; benchmark methodology linked on page; reported comparator settings vary.

        First-party reported result; comparator results retain the source evaluation setup.

        Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        Humanity's Last ExamHard reasoning44.52%HLE no tools

        Effort: Not specified

        contributes to capability
        official board2026-09-09
        LABBench · 2Supporting evidence86.2%
        Reported settings & source

        Selfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Supporting evidence outside the reviewed capability core

        Reviewed 2026-09-06
        LVBench · agenticMultimodal87.8%
        Reported settings & source

        Card labels agentic; linkedmethodology describes only no-tools static setup

        Agentic tool/protocol details unresolved; rawreportedresult only.

        Comparison limit: Card reports agentic result but linked methodology describes only static no-tools protocol.

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Card reports agentic result but linked methodology describes only static no-tools protocol.

        Reviewed 2026-09-06
        LVBench · staticMultimodal87.1%
        Reported settings & source

        No tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024

        Frame budgets differ; table labels Gemini3.8static explicitly.

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        OSWorld partial · 2.0 pre-08.08Agentic59%
        Reported settings & source

        Partialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness

        Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports.

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        lab self-report

        Mixed source task revisions and best-of-three versus provider reporting prevent a uniform common-core unit.

        Reviewed 2026-09-06
        Terminal-Bench · 2.1Coding89.4%
        Reported settings & source

        Terminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        Terminal-Bench · 4.0Coding19.1%
        Reported settings & source

        Officialpublicboard highest scoring thinking level; nativeagents may differ

        Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        Terminal-Bench 4.0 · 4.0Coding19.1%
        Reported settings & source

        Maximum reported across reasoning efforts; OpenAI research environment or API

        Provider-published result; comparator measurements are not automatically independently reproduced.

        GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Gemini 3.8 Flash · reviewed 2026-09-06

        Published configuration

        Effort: Not specified

        contributes to capability
        lab self-reportReviewed 2026-09-06
        ARC-AGI-2Hard reasoning
        GDPval-AAAgentic
        LiveCodeBenchCoding
        LMArena Text ArenaHuman pref
        MMLU-ProKnowledge
        OSWorld-VerifiedAgentic
        SWE-bench ProAgentic
        SWE-bench VerifiedAgentic
        Terminal-Bench 2.1Agentic