RankingGemini 3.8 Flash
Gemini 3.8 Flash
22 published benchmark measures · 7 benchmark families contribute across 4 task areas. 4 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 3 effort levels across 58 benchmark/harness combinations →
Reported effort · High thinking + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- High thinking: 1 observations
- Not specified: 25 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
7 contributing families across 4 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Gemini 3
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://ai.google.dev/gemini-api/docs/models
- Default Capability family coverage
/badge/gemini-3.8-flash.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Artificial Analysis Coding Agent Index v1.4 · v1.4 | Supporting evidence | 61.2 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index v4.1.1 · v4.1.1 | Supporting evidence | 58.7 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Difficult | Supporting evidence | 56.5%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Solvable | Supporting evidence | 88.8%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CharXiv Reasoning | Multimodal | 86.2%Reported settings & sourceNo tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE · 1.1 | Coding | 73.7%Reported settings & sourceDatacurve highest reported effort; Gemini3.8selfcomputed mini-swe-agent highthinking Card Opus74 explicitly corrected as erroneous in linkedmethodology; omitted. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 73.8% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 73.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 61.4%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Extended (score) · 1.1 | Supporting evidence | 56.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Main (score) · 1.1 | Supporting evidence | 43.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDP.PDF | Supporting evidence | 35%Reported settings & sourceAll-pass rate; allmodels selfcomputed byGoogle Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1545Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond | Hard reasoning | 95.253% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 95.3%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey Legal Agent Benchmark · source release snapshot; version not specified | Supporting evidence | 10%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional (length-adjusted) · not specified | Supporting evidence | 52.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE-Verified · source release snapshot; version not specified | Hard reasoning | 54.9%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Humanity's Last Exam | Hard reasoning | 44.52% | HLE no toolscontributes to capability | official board | 2026-09-09 |
| LABBench · 2 | Supporting evidence | 86.2%Reported settings & sourceSelfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LVBench · agentic | Multimodal | 87.8%Reported settings & sourceCard labels agentic; linkedmethodology describes only no-tools static setup Agentic tool/protocol details unresolved; rawreportedresult only. Comparison limit: Card reports agentic result but linked methodology describes only static no-tools protocol. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LVBench · static | Multimodal | 87.1%Reported settings & sourceNo tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 1024 Frame budgets differ; table labels Gemini3.8static explicitly. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 pre-08.08 | Agentic | 59%Reported settings & sourcePartialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 89.4%Reported settings & sourceTerminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 4.0 | Coding | 19.1%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 4.0 · 4.0 | Coding | 19.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Gemini 3.8 Flash · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| LMArena Text Arena | Human pref | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| OSWorld-Verified | Agentic | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |