Show contentsHide contents
- Summary
- At a glance
- Who can use Argon today
- Google's benchmark claims
- What independent tests show
- Why the scores disagree
- What it really costs
- The 1M output token limit
- Cybersecurity: the headline use case
- What Google says it uses Argon for internally
- Safety approach
- Where Argon is weaker
- Which model for which job
- What we don't know yet
- Frequently asked questions
- Sources
Gemini 4 Argon is first on the Vals Index, an independent benchmark of finance, coding, legal and tax work, at 68.90%.1 On the Artificial Analysis Intelligence Index, a different independent benchmark, it ties for third at 53 points, behind two Claude models.23 Both results are true, and the gap between them is the best summary of Google's new frontier model: competitive at the top, leading on some kinds of work, and not the outright winner Google's own benchmark table suggests.
Most people can't use it yet. Google announced Argon on 30 September 2026 and is releasing it first to trusted cyber defenders through its Fairwind Program. Developers, businesses and consumers are to follow "as soon as possible", with no date given.4
This article covers everything published about Argon so far: Google's claims, the independent results that confirm or complicate them, what it really costs, who can get it, and where it falls short. Every Google figure is marked as vendor-reported. Last checked 1 October 2026.
Summary
- Access: limited to approved partners in Google's Fairwind Program for now. Paid API customers and Google AI Ultra subscribers are next, with no date announced.4
- Price: USD 2 per million input tokens and USD 10 per million output tokens, described by Google as an introductory price. Cached input is 95% cheaper.4 Vals lists USD 4 and USD 20, which The Decoder reports as the regular price.53
- Independent rankings: first on the Vals Index.1 Level with GPT-6 Astra and Claude Fable 5.1 on the Artificial Analysis Intelligence Index, five points behind Claude Opus 5.5.3
- Google's own table: against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, Argon is best or joint best on 14 of 19 rows and loses five, mostly in terminal-based coding and science tasks.4
- Cost per task: cheap per token, but verbose. Artificial Analysis found it uses more than twice as many output tokens per task as GPT-6 Astra.3
- Output limit: 1 million tokens, according to Google.4 Vals lists 262K.5
At a glance
| Gemini 4 Argon | |
|---|---|
| Announced | 30 September 20264 |
| Who can use it now | Approved Fairwind Program partners6 |
| Next in line | Paid API customers and Google AI Ultra subscribers, date not announced4 |
| Introductory price | USD 2 input / USD 10 output per million tokens. Cached input 95% off.4 |
| Listed standard price | USD 4 / USD 20 per million tokens, per Vals and The Decoder53 |
| Context window (input) | 1 million tokens52 |
| Output limit | 1 million tokens according to Google, up from 64K.4 Vals lists 262K.5 |
| Inputs | Text, images, video and audio, per The Decoder.3 Artificial Analysis lists text and image.2 |
| Output | Text2 |
| Weights | Proprietary5 |
Who can use Argon today
Argon is going first to the Fairwind Program, which Google describes as giving "high-priority defenders (like governments, healthcare providers, and telecommunications services)" early access to advanced models.6 Google says it works with more than 650 partners, though not all of them get Argon. The page describes Argon access as exclusive to "a set of" Fairwind partners.6
The terms are strict:6
- Only internal cybersecurity, incident response or penetration-testing teams may be given access, and organisations must track employee access and use.
- Partners must use user-level authentication and phishing-resistant MFA.
- Access may not be shared, redistributed or sold.
- Google runs background checks on applicants' security history and record of ethical operations.
- Permitted dual-use work is limited to authorised threat simulation, reverse engineering and malware analysis, for defensive and academic purposes.
For these trusted defenders and for Google's own teams, Argon is released without cyber guardrails, so they can use its full security capability.4 Google also says it is taking part in the US government's voluntary pre-release access process while it gradually widens access.4
Google's benchmark claims
Google published a table comparing Argon with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and Claude Opus 5.5, with methodology in a separate evaluation document.47 All figures below are as reported by Google.
| Area | Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% | |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% | |
| Harvey's Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% | |
| Agentic coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% | |
| Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% | |
| Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% | |
| ML engineering | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Science and maths | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% | |
| RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% | |
| Long context | GraphWalks, up to 128K | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks, 256K to 1M | 84.2% | 71.8% | 65.0% | 66.8% | |
| Computer use | Agent's Last Exam | 39.5% | 34.2% | n/a | 38.2% |
| OSWorld-2.0 | 69.2% | 72.6% | n/a | n/a | |
| Multimodal | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench | 91.7% | 87.5% | 79.7% | 83.7% | |
| Cybersecurity | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Source: Google, Gemini 4 Argon announcement and evaluation methodology. Bold marks the best score in each row. n/a means Google did not report a score.47
Three things in this table deserve attention.
Argon loses five rows. GPT-6 Astra is ahead on FrontierSWE v2 by 10.5 points, on Terminal-Bench Science by 10.5 points, and on OSWorld-2.0. Claude Opus 5.5 is ahead on Terminal-bench 4.0 by 9 points and on PostTrainBench by 4 points. The losses cluster in long, terminal-driven engineering and science work.
The biggest leads are in knowledge work. On Harvey's Legal Agent Benchmark, Argon's 19.6% is roughly three times the next model's score. The gaps on Finance Agent v2 and on the long-context half of GraphWalks are also large.
A strong rival is missing. Google compared Argon with Claude Opus 5.5 and Fable 5.1, but not with Claude Sonnet 5.5, which is second on the same Vals Index Google cites, ahead of both of them.1
What independent tests show
Three independent evaluators have published Argon results. They broadly confirm that Argon belongs in the top group. They do not confirm a clear lead. For how Claude Opus 5.5 and GPT-6 Astra compare with each other, see our Claude Opus 5.5 vs GPT-6 analysis.
Free to reuse under CC BY 4.0 with a link to this article. Download
Vals Index: first
Vals runs its own GDP-weighted index of finance, coding, legal and tax benchmarks, most of them private.1 It confirms Google's headline figure:
| Rank | Model | Vals Index | Cost per test |
|---|---|---|---|
| 1 | Gemini 4 Argon | 68.90% | USD 15.68 |
| 2 | Claude Sonnet 5.5 | 67.04% | USD 21.34 |
| 3 | Claude Opus 5.5 | 66.97% | USD 32.14 |
| 4 | Claude Fable 5.1 | 65.83% | USD 28.71 |
| 6 | GPT-6 Astra | 63.13% | USD 18.46 |
Source: Vals Index v2.1, updated 30 September 2026. Vals reports a standard error of about ±0.97 points for Argon's score.15
The Decoder counts top-five finishes on 20 of the 22 Vals benchmarks Argon was tested on.3 It is first on Finance Agent v2, tied first with GPT-6 Astra on IOI, and second on Vibe Code Bench, Code Migration, Tax Agent Bench and CyberBench.5 Vals ran it at "high" reasoning effort with temperature 1.0.5
One caveat on the cost column: Vals calculates cost per test at USD 4 and USD 20 per million tokens, not at Google's introductory price.5
Artificial Analysis: tied third
The Artificial Analysis Intelligence Index combines ten evaluations, including agentic knowledge work, Terminal-Bench 4.0, Humanity's Last Exam and long-context reasoning.2 At Argon's highest reasoning level, "High", it scores 53.2
According to The Decoder's reading of the Artificial Analysis results:3
| Model | Intelligence Index |
|---|---|
| Claude Opus 5.5 | 58 |
| Claude Sonnet 5.5 | 56 |
| Gemini 4 Argon | 53 |
| GPT-6 Astra | 53 |
| Claude Fable 5.1 | 53 |
| GPT-6.1 Sol | 52 |
That is a 23-point jump over Google's previous frontier model, Gemini 3.1 Pro Preview.3 Two findings from the same testing stand out:
- Agentic work improved most. Argon takes first place on AutomationBench-AA at 77.5%. On Terminal-Bench 4.0 it scores 57%, behind Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%) and GPT-6 Astra (59%).3
- It hallucinates far less, but knows less. On AA-Omniscience, Argon's hallucination rate is 15%, against 51% for GPT-6 Astra. Its accuracy is 50% against Astra's 63%, and the two finish roughly level on the overall score, at 42 and 43.3
Human preference: first for text
On Arena, where people compare anonymous model outputs head to head, Argon (High) is reported first in the Text Arena with 1,525 points, 20 ahead of second place. It ranks eighth in Code Arena: WebDev.3 These figures come from Arena's announcement as reported by The Decoder. We have not checked the leaderboard directly.
Why the scores disagree
The same model, tested at around the same time, ranks first on one index and third on another. Three reasons explain most of it:
- Different task mixes. Vals weights finance most heavily, in proportion to its share of US GDP, and Argon is strongest there.1 Artificial Analysis gives more weight to terminal-based coding and broad reasoning, where Claude leads.2
- Different harnesses. On Terminal-bench 4.0, Google's table shows Claude Opus 5.5 at 66.4%.4 Anthropic says it reports Opus 5.5's Terminal-Bench 4.0 result at the model's highest effort setting, "xhigh".8 Artificial Analysis measured Opus 5.5 at 60% on the same benchmark.3 The scaffolding around a model can move results by several points.
- Safeguards change scores. Anthropic evaluates Claude Opus 5.5 and Fable 5.1 with their production safeguards switched on. When those intervene, cybersecurity tasks are completed by the older Claude Opus 4.8, and biology tasks by Claude Opus 5. Anthropic says this likely reduces its models' scores.89 Argon's security results come from a model that trusted defenders get without cyber guardrails.4
What it really costs
Free to reuse under CC BY 4.0 with a link to this article. Download
| Model | Input (per M tokens) | Output (per M tokens) | Cached input |
|---|---|---|---|
| Gemini 4 Argon (introductory) | USD 2 | USD 10 | 95% off input4 |
| Gemini 4 Argon (standard, as listed by Vals) | USD 4 | USD 20 | not stated5 |
| Claude Opus 5.5 | USD 4 | USD 20 | USD 0.208 |
| Claude Fable 5.1 | USD 10 | USD 50 | USD 0.259 |
| GPT-6 Astra | USD 10 | USD 50 | USD 1.0010 |
| GPT-6.1 Sol | USD 2 | USD 10 | USD 0.1010 |
Per token, Argon at its introductory price is among the cheapest frontier models. Per task, the picture changes, because Argon writes a lot:
- Artificial Analysis found Argon uses an average of 62,000 output tokens per task, against 27,000 for GPT-6 Astra.3
- At the introductory price, an Intelligence Index task costs USD 1.99, about 60% of GPT-6 Astra's USD 3.26.23
- At the standard price, The Decoder calculates USD 3.98 per task, about 20% more than Astra.3
On the Vals Index, which already uses the standard price, Argon is cheaper per test than every Claude model ranked near it.1
The 1M output token limit
Google's most unusual change is the output limit: 1 million tokens, up from 64,000, which Google calls industry-leading. The idea is to let the model "think deeply and generate hundreds of thousands of tokens in a single trajectory".4 The Decoder reports a new Long Decode Continuation feature in the Gemini API that pauses very long responses and resumes them through follow-up requests, so reasoning does not time out.3
Two things to note:
- The limit is not consistent across sources. Vals' model page lists a maximum output of 262K tokens.5 We could not resolve the difference. It may reflect what a particular endpoint or access tier exposes.
- A full-length answer is not cheap. By our arithmetic, one million output tokens costs USD 10 at the introductory price and USD 20 at the listed standard price, before any input or reasoning overhead.
Cybersecurity: the headline use case
Google trained Argon specifically for defensive security and says it can "autonomously find, validate, and patch critical software vulnerabilities".4
- CWE-bench v1 (fixing security vulnerabilities): Argon ties for first at 68% with GPT-6 Astra and Grok 4.7. Claude Opus 5.5 scores 67% and Claude Fable 5.1 58%.46
- Vals CyberBench: second at 77.86%, 0.12 points behind GPT-6 Sol.5
- Real-world claim: Google says Wiz, using Argon through its Scan for Good programme, found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, which previous frontier models had missed.4 This is Google's account, and we have not seen the vulnerability disclosed independently.
Fairwind partners can use Argon inside CodeMender, Google's code security agent for automated fixes.6
What Google says it uses Argon for internally
Google lists several internal results. These are Google's own claims, and none has been independently verified:4
- Memory: a team of Argon agents analysed fleet-wide profiling data and applied memory optimisations across Google's data centres, freeing more than 300 TiB, with an estimated 500 TiB to 1 PiB in total savings.
- C and C++ to Rust: Argon agents are migrating codebases from tens of thousands of lines up to the 800K+ line Fuchsia Zircon kernel, with automated and manual review before production.
- Video decoding: for libgav1, Argon replaced 32K lines of SIMD code in an existing Rust port. Google says the result is a memory-safe decoder that runs 2.7 times faster than that port, with identical output.
- Quantum research: in one example, Argon beat a published baseline for quantum circuit resources by 40% "in a matter of minutes".
Safety approach
Google says it is strengthening safeguards in four areas before broad availability:4
- Misuse: refusing harmful cyber and CBRN requests while preserving legitimate dual-use research, including monitoring the model's internal activations to spot misuse.
- Prompt injection: Google calls Argon its most resilient model yet and says it leads on Gray Swan's Indirect Prompt Injection benchmark.
- Misalignment: monitoring Argon's chain of thought and actions, and stopping execution when needed.
- Hardened environments: isolating and sealing sandboxes before high-risk training and evaluation.
Google also urges the rest of the industry to "preserve reasoning transparency".4 That is a pointed contrast with OpenAI's GPT-6 Astra, which TechCrunch described as possibly OpenAI's most controversial model because of a reasoning technique that obscures its chain of thought.11
Where Argon is weaker
From the independent results so far:
- Terminal-based coding: fifth on Vals' Terminal-Bench 4.0, at 57.58%.5 Behind Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Astra in Artificial Analysis testing.3
- Computer use: 4.83% on Vals' CUA-bench, seventh of eight models, at a cost of USD 193.78 per test.5 GPT-6 Astra is ahead on OSWorld-2.0 even in Google's own table.4
- Medical scribing: 15th on Vals' MedScribe.5
- Web development: eighth on Arena's WebDev leaderboard, as reported by The Decoder.3
- Raw knowledge recall: lower accuracy than GPT-6 Astra on AA-Omniscience.3
Which model for which job
What we don't know yet
- When Argon will be generally available. Google says "as soon as possible".4
- When the introductory price ends, and the official standard price. The USD 4 / USD 20 figure comes from Vals and The Decoder, not from Google's announcement.53
- The real output limit on each endpoint: 1M according to Google, 262K according to Vals.45
- How Argon performs outside cyber defence for ordinary customers, since nearly all testing so far is by Google, its partners and benchmark operators.
We will update this article when Google opens wider access or publishes official standard pricing.
Frequently asked questions
- When was Gemini 4 Argon released?
- Google announced Gemini 4 Argon on 30 September 2026. It is rolling out first to trusted cyber defenders through Google's Fairwind Program. Google says developers, enterprises and consumers will follow, starting with paid API customers and Google AI Ultra subscribers, but it has not given a date.
- How much does Gemini 4 Argon cost?
- Google lists an introductory API price of USD 2 per million input tokens and USD 10 per million output tokens, with cached input 95% cheaper. Vals AI lists the model at USD 4 and USD 20, and The Decoder reports that the regular price will rise to that level after the introductory period. Google has not said when the introductory price ends.
- Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?
- It depends on the test. Argon leads the independent Vals Index at 68.90%, ahead of Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Astra. On the Artificial Analysis Intelligence Index it scores 53, level with GPT-6 Astra and Claude Fable 5.1 and behind Claude Opus 5.5 at 58. In Google's own table it is best or joint best on 14 of 19 rows and loses five.
- What is the Gemini 4 Argon output token limit?
- Google says Argon's output limit is 1 million tokens, up from 64,000 in its previous models. Vals AI's model page lists a maximum output of 262K tokens. We could not resolve the difference at the time of writing; it may reflect what is exposed on a given endpoint or tier.
- Can I use Gemini 4 Argon now?
- Only if your organisation is an accepted partner in Google's Fairwind Program, which is aimed at high-priority defenders such as governments, healthcare providers and telecommunications services. Everyone else has to wait for the wider rollout.
- Is Gemini 4 Argon good for cybersecurity?
- Google trained it specifically for defensive security work. On CWE-bench v1 it ties for first at 68% with GPT-6 Astra and Grok 4.7, and on Vals' CyberBench it places second, 0.12 points behind GPT-6 Sol. Comparisons with Claude are complicated because Anthropic evaluates its models with safeguards that hand some security tasks to older models.
Sources
-
Vals AI, "Vals Index", version 2.1, updated 30 September 2026. https://www.vals.ai/benchmarks/vals_index ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7
-
Artificial Analysis, "Gemini 4 Argon (High) Intelligence, Performance & Price Analysis". https://artificialanalysis.ai/models/gemini-4-argon ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
Matthias Bastian, "Google Gemini 4 Argon closes the gap with OpenAI and Anthropic but doesn't take a clear lead", The Decoder, 1 October 2026. https://the-decoder.com/google-gemini-4-argon-closes-the-gap-with-openai-and-anthropic-but-doesnt-take-a-clear-lead/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21
-
Koray Kavukcuoglu, "Gemini 4 Argon: our next era of frontier intelligence", Google, 30 September 2026. Includes Google's benchmark comparison table. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21 ↩22 ↩23 ↩24 ↩25 ↩26
-
Vals AI, "Gemini 4 Argon" model page, released 30 September 2026. https://www.vals.ai/models/google_gemini-4-argon ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18
-
Google DeepMind, "Fairwind Program". https://deepmind.google/fairwind-program/ ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
Google DeepMind, Gemini 4 Argon model evaluation methodology (PDF). https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf ↩ ↩2
-
Anthropic, "Claude Opus" model page, including Opus 5.5 pricing and benchmark notes. https://www.anthropic.com/claude/opus ↩ ↩2 ↩3
-
Anthropic, "Claude Fable" model page, including Fable 5.1 pricing and safeguards. https://www.anthropic.com/claude/fable ↩ ↩2
-
OpenAI, "API Pricing", standard processing for context lengths under 272K. https://openai.com/api/pricing/ ↩ ↩2
-
Lucas Ropek, "OpenAI launches Astra, its powerful (and controversial) new model", TechCrunch, 3 September 2026. https://techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/ ↩