RankingGemini 4 Argon

Gemini 4 Argon

Compare models

Data updated 30 Sept 2026

18 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Gemini 4 Argon capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

AgenticNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 13/13 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    Hard reasoningNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      CodingNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        Human prefNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          KnowledgeNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            MultimodalNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Long contextNo comparable evidenceUnknown

              0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

              Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

              Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

              Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

              Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

                Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

                Reported effort · Not specified

                Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

                Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

                Inspect each result and its source ↓ · Download effort evidence
                Score contributions and missing evidence

                0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.

                Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

                Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

                Model information & shareable badge
                Lab
                Google
                Catalog status
                limited
                Availability
                Trusted cyber defenders through the Fairwind Program. Broader developer, enterprise and consumer access is planned, starting with paid API customers and Google AI Ultra subscribers.
                Family
                Gemini 4
                Released
                2026-09-30
                Context
                —
                API list price
                $2 input / $10 output per million tokens
                Announced introductory API pricing, not general availability. Cache reads 95% below input price. After the introductory period, $4 input and $20 output per million tokens; expiry date not published.
                License
                proprietary
                Model card
                https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
                Default Capability family coverage
                Documented-evidence family coverage/badge/gemini-4-argon.svg
                Benchmark scores & sources

                Original results, evaluation harnesses, and evidence behind this model.

                BenchmarkBucketScoreHarnessEvidenceSource-recorded date
                Agent's Last ExamSupporting evidence39.5%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed in default ALE-Claw harness with a five-hour window and safety filters. Flagged responses return empty strings; episode may continue. Astra and Opus from the official board. Binary pass rate. Fable absent.

                Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed in default ALE-Claw harness with a five-hour window and safety filters. Flagged responses return empty strings; episode may continue. Astra and Opus from the official board. Binary pass rate. Fable absent. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 15: Agent's Last Exam / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                AutomationBenchAgentic51.3%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Private test set; results sourced from the official Zapier public leaderboard.

                Provider-published launch claim, not independent measurement by Google for every row. Private test set; results sourced from the official Zapier public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 2: AutomationBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                AutomationBench-AASupporting evidence77.51346945899195%Artificial Analysis AutomationBench-AA

                Effort: Not specified

                official board

                Supporting evidence outside the reviewed capability core

                2026-09-30
                ChartographyMultimodal71.6%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Without tools; results from the official Surge public leaderboard.

                Provider-published launch claim, not independent measurement by Google for every row. Without tools; results from the official Surge public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 17: Chartography / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                CWE-bench · 1Supporting evidence68%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Official public leaderboard; ranking uses pass@1 with pass@4 tiebreaks. Grid records pass@1 only.

                Provider-published launch claim, not independent measurement by Google for every row. Official public leaderboard; ranking uses pass@1 with pass@4 tiebreaks. Grid records pass@1 only. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 19: CWE-bench 1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                DeepSWE · 1.1Coding77.9%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed with mini-swe agent. Astra from the official public leaderboard; Fable and Opus from their system cards. Highest scoring thinking level per Datacurve.

                Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed with mini-swe agent. Astra from the official public leaderboard; Fable and Opus from their system cards. Highest scoring thinking level per Datacurve. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 5: DeepSWE 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                FrontierSWE · 2Supporting evidence55%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Proximal's official public leaderboard.

                Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Proximal's official public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 6: FrontierSWE 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                GraphWalks · 256K to 1M BFS F1Supporting evidence84.2%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed. 200 problems with context length between 256K and 1M tokens. BFS F1, not pass rate.

                Provider-published launch claim, not independent measurement by Google for every row. All models self-computed. 200 problems with context length between 256K and 1M tokens. BFS F1, not pass rate. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 14: GraphWalks 256K to 1M BFS F1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                GraphWalks · up to 128K BFS F1Supporting evidence99.7%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed. 650 items with context length up to 128K tokens. BFS F1, not pass rate.

                Provider-published launch claim, not independent measurement by Google for every row. All models self-computed. 650 items with context length up to 128K tokens. BFS F1, not pass rate. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 13: GraphWalks up to 128K BFS F1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Harvey's Legal Agent Benchmark · 1Supporting evidence19.6%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Vals AI.

                Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                LABBench · 2Supporting evidence88.8%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed with Linux terminal, pre-installed bioinformatics tools, Python, R and internet access.

                Provider-published launch claim, not independent measurement by Google for every row. All models self-computed with Linux terminal, pre-installed bioinformatics tools, Python, R and internet access. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 11: LABBench 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                LVBenchMultimodal91.7%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison.

                Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                OSWorld · 2.0 offline partial scoreAgentic69.2%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed, maximum of three single-attempt runs; offline subset partial score. Official Docker/evaluator, 1080p, 500 steps, Gemini CUA harness, parallel batch tools, compaction, pyautogui actuation, screenshot-only observations, UI-specific function declarations, safety filters, official 08.08 patch. Astra from its official blog. Anthropic combined online/offline results are excluded.

                Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed, maximum of three single-attempt runs; offline subset partial score. Official Docker/evaluator, 1080p, 500 steps, Gemini CUA harness, parallel batch tools, compaction, pyautogui actuation, screenshot-only observations, UI-specific function declarations, safety filters, official 08.08 patch. Astra from its official blog. Anthropic combined online/offline results are excluded. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 16: OSWorld 2.0 offline partial score / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                PostTrainBench · 1.1Supporting evidence45.3%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. All models self-computed; OpenCode harness; 10-hour budget on one NVIDIA H100 GPU. Weighted aggregate across four base models and seven benchmarks.

                Provider-published launch claim, not independent measurement by Google for every row. All models self-computed; OpenCode harness; 10-hour budget on one NVIDIA H100 GPU. Weighted aggregate across four base models and seven benchmarks. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 9: PostTrainBench 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                RiemannBenchSupporting evidence76%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Results sourced from the official Surge public leaderboard.

                Provider-published launch claim, not independent measurement by Google for every row. Results sourced from the official Surge public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 12: RiemannBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Terminal-Bench · 4.0Coding57.4%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed; other models from the official public leaderboard. Highest scoring thinking level reported by Terminal-Bench authors. Argon agent identity is not specified in this methodology.

                Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed; other models from the official public leaderboard. Highest scoring thinking level reported by Terminal-Bench authors. Argon agent identity is not specified in this methodology. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 8: Terminal-Bench 4.0 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Terminal-Bench-Science · 0.1Coding57.6%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed with 6x verifier timeout to address verification timeout issues. Other models from the official public leaderboard. Modified verifier timeout is excluded from matched comparison.

                Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed with 6x verifier timeout to address verification timeout issues. Other models from the official public leaderboard. Modified verifier timeout is excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 10: Terminal-Bench Science 0.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Vals Finance Agent · 2Supporting evidence65.4%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Vals AI.

                Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 3: Vals Finance Agent 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Vals IndexSupporting evidence68.9%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Vals AI composite, weighted by contribution to US GDP; overlapping component benchmarks do not add an aggregate ranking vote.

                Provider-published launch claim, not independent measurement by Google for every row. Vals AI composite, weighted by contribution to US GDP; overlapping component benchmarks do not add an aggregate ranking vote. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 1: Vals Index / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                Vibe Code Bench · 1.1Supporting evidence91.9%
                Reported settings & source

                Gemini API highest thinking settings; exact named effort not specified by Google. Results sourced from the official Vals AI public leaderboard.

                Provider-published launch claim, not independent measurement by Google for every row. Results sourced from the official Vals AI public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date.

                Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 7: Vibe Code Bench 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30

                Published configuration

                Effort: Not specified

                lab self-report

                Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison.

                Reviewed 2026-09-30
                ARC-AGI-2Hard reasoning————
                DeepSWE v1.1Agentic————
                GDPval-AAAgentic————
                GPQA DiamondHard reasoning————
                Humanity's Last ExamHard reasoning————