RankingGemini 4 Argon
Gemini 4 Argon
18 published benchmark measures · 0 benchmark families contribute across 0 task areas. 0 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 20 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
0 contributing families across 0 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Catalog status
- limited
- Availability
- Trusted cyber defenders through the Fairwind Program. Broader developer, enterprise and consumer access is planned, starting with paid API customers and Google AI Ultra subscribers.
- Family
- Gemini 4
- Released
- 2026-09-30
- Context
- —
- API list price
- $2 input / $10 output per million tokens
- License
- proprietary
- Model card
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
- Default Capability family coverage
/badge/gemini-4-argon.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| Agent's Last Exam | Supporting evidence | 39.5%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed in default ALE-Claw harness with a five-hour window and safety filters. Flagged responses return empty strings; episode may continue. Astra and Opus from the official board. Binary pass rate. Fable absent. Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed in default ALE-Claw harness with a five-hour window and safety filters. Flagged responses return empty strings; episode may continue. Astra and Opus from the official board. Binary pass rate. Fable absent. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 15: Agent's Last Exam / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| AutomationBench | Agentic | 51.3%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Private test set; results sourced from the official Zapier public leaderboard. Provider-published launch claim, not independent measurement by Google for every row. Private test set; results sourced from the official Zapier public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 2: AutomationBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| AutomationBench-AA | Supporting evidence | 77.51346945899195% | Artificial Analysis AutomationBench-AA | official board | 2026-09-30 |
| Chartography | Multimodal | 71.6%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Without tools; results from the official Surge public leaderboard. Provider-published launch claim, not independent measurement by Google for every row. Without tools; results from the official Surge public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 17: Chartography / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| CWE-bench · 1 | Supporting evidence | 68%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Official public leaderboard; ranking uses pass@1 with pass@4 tiebreaks. Grid records pass@1 only. Provider-published launch claim, not independent measurement by Google for every row. Official public leaderboard; ranking uses pass@1 with pass@4 tiebreaks. Grid records pass@1 only. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 19: CWE-bench 1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| DeepSWE · 1.1 | Coding | 77.9%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed with mini-swe agent. Astra from the official public leaderboard; Fable and Opus from their system cards. Highest scoring thinking level per Datacurve. Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed with mini-swe agent. Astra from the official public leaderboard; Fable and Opus from their system cards. Highest scoring thinking level per Datacurve. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 5: DeepSWE 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| FrontierSWE · 2 | Supporting evidence | 55%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Proximal's official public leaderboard. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Proximal's official public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 6: FrontierSWE 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| GraphWalks · 256K to 1M BFS F1 | Supporting evidence | 84.2%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. All models self-computed. 200 problems with context length between 256K and 1M tokens. BFS F1, not pass rate. Provider-published launch claim, not independent measurement by Google for every row. All models self-computed. 200 problems with context length between 256K and 1M tokens. BFS F1, not pass rate. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 14: GraphWalks 256K to 1M BFS F1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| GraphWalks · up to 128K BFS F1 | Supporting evidence | 99.7%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. All models self-computed. 650 items with context length up to 128K tokens. BFS F1, not pass rate. Provider-published launch claim, not independent measurement by Google for every row. All models self-computed. 650 items with context length up to 128K tokens. BFS F1, not pass rate. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 13: GraphWalks up to 128K BFS F1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Harvey's Legal Agent Benchmark · 1 | Supporting evidence | 19.6%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 4: Harvey's Legal Agent Benchmark 1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| LABBench · 2 | Supporting evidence | 88.8%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. All models self-computed with Linux terminal, pre-installed bioinformatics tools, Python, R and internet access. Provider-published launch claim, not independent measurement by Google for every row. All models self-computed with Linux terminal, pre-installed bioinformatics tools, Python, R and internet access. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 11: LABBench 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| LVBench | Multimodal | 91.7%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Provider-published launch claim, not independent measurement by Google for every row. All models self-computed without tools. Gemini 1 FPS; Astra 800 frames, Fable 300 frames and Opus 600 frames due to API limits. Unequal frame budgets are excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 18: LVBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| OSWorld · 2.0 offline partial score | Agentic | 69.2%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed, maximum of three single-attempt runs; offline subset partial score. Official Docker/evaluator, 1080p, 500 steps, Gemini CUA harness, parallel batch tools, compaction, pyautogui actuation, screenshot-only observations, UI-specific function declarations, safety filters, official 08.08 patch. Astra from its official blog. Anthropic combined online/offline results are excluded. Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed, maximum of three single-attempt runs; offline subset partial score. Official Docker/evaluator, 1080p, 500 steps, Gemini CUA harness, parallel batch tools, compaction, pyautogui actuation, screenshot-only observations, UI-specific function declarations, safety filters, official 08.08 patch. Astra from its official blog. Anthropic combined online/offline results are excluded. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 16: OSWorld 2.0 offline partial score / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| PostTrainBench · 1.1 | Supporting evidence | 45.3%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. All models self-computed; OpenCode harness; 10-hour budget on one NVIDIA H100 GPU. Weighted aggregate across four base models and seven benchmarks. Provider-published launch claim, not independent measurement by Google for every row. All models self-computed; OpenCode harness; 10-hour budget on one NVIDIA H100 GPU. Weighted aggregate across four base models and seven benchmarks. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 9: PostTrainBench 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| RiemannBench | Supporting evidence | 76%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from the official Surge public leaderboard. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from the official Surge public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 12: RiemannBench / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Terminal-Bench · 4.0 | Coding | 57.4%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed; other models from the official public leaderboard. Highest scoring thinking level reported by Terminal-Bench authors. Argon agent identity is not specified in this methodology. Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed; other models from the official public leaderboard. Highest scoring thinking level reported by Terminal-Bench authors. Argon agent identity is not specified in this methodology. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 8: Terminal-Bench 4.0 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Terminal-Bench-Science · 0.1 | Coding | 57.6%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Argon self-computed with 6x verifier timeout to address verification timeout issues. Other models from the official public leaderboard. Modified verifier timeout is excluded from matched comparison. Provider-published launch claim, not independent measurement by Google for every row. Argon self-computed with 6x verifier timeout to address verification timeout issues. Other models from the official public leaderboard. Modified verifier timeout is excluded from matched comparison. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 10: Terminal-Bench Science 0.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Vals Finance Agent · 2 | Supporting evidence | 65.4%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from Vals AI. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from Vals AI. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 3: Vals Finance Agent 2 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Vals Index | Supporting evidence | 68.9%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Vals AI composite, weighted by contribution to US GDP; overlapping component benchmarks do not add an aggregate ranking vote. Provider-published launch claim, not independent measurement by Google for every row. Vals AI composite, weighted by contribution to US GDP; overlapping component benchmarks do not add an aggregate ranking vote. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 1: Vals Index / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| Vibe Code Bench · 1.1 | Supporting evidence | 91.9%Reported settings & sourceGemini API highest thinking settings; exact named effort not specified by Google. Results sourced from the official Vals AI public leaderboard. Provider-published launch claim, not independent measurement by Google for every row. Results sourced from the official Vals AI public leaderboard. Methodology https://deepmind.google/models/evals-methodology/gemini-4-argon, SHA256 ff1df6bdeddc4c0f48840c09a8a7813b10ae7ad9e053892da68dabf64f09bff7. Preserve source rounding. Methodology says capabilities as of September 2026 and results as of October 2026; review date is source retrieval, not evaluation date. Comparison limit: Provider-published claim. Independent source results are admitted separately with exact checkpoint, effort and compatible protocol; this grid does not establish a matched comparison. Google Gemini 4 Argon launch and evaluation methodology · Launch grid row 7: Vibe Code Bench 1.1 / Gemini 4 Argon; methodology Additional Details · reviewed 2026-09-30 | Published configuration | lab self-report | Reviewed 2026-09-30 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |