TL;DR
- Google DeepMind announced Gemini 4 Argon on September 30 — its first frontier model above the 3.8 Flash line since February. The model is currently accessible only to trusted cyber defenders and government partners through the Fairwind Program (650+ participating organizations). - On DeepSWE v1.1 — the benchmark closest to real-world software engineering — Argon scores 77.9%, ahead of GPT-6 Astra (74.1%) and Claude Opus 5.5 (74.2%). On Vals Index, it ranks #1 at 68.9%, using approximately one quarter of Sonnet 5.5's output tokens per task. - The output token limit jumps from 64K to 1M via Long Decode Continuation — a new API feature that pauses long responses and resumes them across calls. The native output limit is 262K tokens; the 1M figure is achieved through continuation. This is the most consequential engineering change in the release. - Introductory pricing: $2 input / $10 output per million tokens (same as GPT-6.1 Sol at DevDay). Standard pricing after the introductory period: $4/$20. Cached input: 95% discount (same structure as Step 5 Preview). - Argon is being released without cyber guardrails to Fairwind participants and Google's internal teams — giving them full offensive-adjacent capability for defensive work: autonomously finding, validating, and patching vulnerabilities. Wiz used Argon to find a critical vulnerability in healthcare software that previous frontier models had missed. - On the Gray Swan Indirect Prompt Injection benchmark, Argon leads all frontier models in robustness. - Broad access — paid API customers and Google AI Ultra subscribers — is coming "as soon as possible." No date given.
The way Google released Gemini 4 Argon is as significant as what it released.
Every major frontier model launch in the past 18 months has followed the same pattern: announce, release to API, roll out to consumers, add enterprise controls. Argon reverses this. Cybersecurity teams and government partners get the model first — specifically without the safety guardrails that would constrain its offensive-adjacent capabilities — while everyone else waits.
This is not a soft launch strategy. It is a deliberate architecture for who gets frontier capability and when. And the reasoning is coherent: a model capable of autonomously finding and exploiting zero-days needs different controls for a defender than for a general-purpose API customer. Google is building the trust infrastructure before it builds the distribution.
What Argon actually is
Gemini 4 Argon is positioned as Google's answer to GPT-6 Astra and Claude Fable 5.1 — the frontier tier, not the fast tier. It was announced by Koray Kavukcuoglu, SVP of Google DeepMind, as built for "deep reasoning across complex, long-horizon workflows" in software engineering, legal and financial knowledge work, and cybersecurity defense.
The benchmark that matters most for enterprise teams is DeepSWE v1.1 — an evaluation of real-world software engineering capability, not synthetic coding puzzles. Argon scores 77.9%, establishing a new state of the art. GPT-6 Astra sits at 74.1%; Claude Opus 5.5 at 74.2%. On Vals Index — which evaluates real-world task completion across enterprise workflows — Argon ranks first at 68.9%, at an average cost of $15.68 per task.
The 77.9% on DeepSWE is vendor-reported. No third party has independently reproduced Argon's benchmarks as of publication. Treat them as Google's published claims pending independent evaluation.
The 1M output token claim
The headline engineering change is the output token limit, which Google describes as 1M tokens — up from 64K on the previous generation.
This requires a clarification that most coverage has glossed over. The native output limit is 262K tokens. The 1M figure is achieved through Long Decode Continuation — a new API feature that pauses long responses and resumes them across multiple API calls, stitching them together into a single coherent trajectory. Artificial Analysis confirmed reaching 1M output tokens through this mechanism.
The distinction matters for workflow design. A true 1M output limit in a single call would allow a model to write an entire codebase migration, complete a full financial analysis, or produce a legal draft end-to-end without stopping. Long Decode Continuation achieves a similar functional result but requires the calling infrastructure to handle the pause-and-resume protocol — which adds engineering overhead that a raw 1M output limit would not.
The capability is real. The architecture is more complex than the headline implies. If you design around "1M output tokens," confirm your infrastructure handles Long Decode Continuation before building production dependencies on it.
On the reasoning side: Google notes that "when the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go." The practical test is whether the reasoning quality at very long outputs degrades — a question that requires external evaluation rather than Google's own assessment.
The cybersecurity architecture decision
This is the part of the Argon launch that deserves most attention.
Google is releasing Argon without cyber guardrails to two groups: internal Google teams, and Fairwind Program participants. Fairwind launched September 2 and includes more than 650 organizations — government agencies, critical infrastructure operators in healthcare, telecom, energy, and finance, and security partners including CrowdStrike, Palo Alto Networks, Snowflake, and Wiz.
Without cyber guardrails means Fairwind participants can use Argon's full capability for offensive-adjacent defensive work: autonomously finding vulnerabilities, validating exploitability, and generating proof-of-concept evidence. Wiz demonstrated this already through its Scan for Good program, where Argon surfaced a critical vulnerability exposing sensitive personal data across healthcare software used by hospitals worldwide — a vulnerability that previous frontier models missed.
On CWE-bench v1 (the benchmark for vulnerability remediation), Argon ties for first with a score of 68%. Google is more emphatic about the discovery side: on Wiz's black-box penetration testing benchmark, Argon outperforms Gemini 3.8 Flash Cyber across attack surface mapping, vulnerability identification, and proof-of-concept generation, across codebases spanning 20 programming languages.
The decision to release without guardrails to trusted defenders is not unusual in security tooling — penetration testing tools have always worked this way. What is new is applying this architecture to a frontier LLM capable of autonomous multi-step reasoning across an entire attack surface. The Fairwind gatekeeping mechanism — vetting organizations before access, not just after — is Google's attempt to build the trust infrastructure that makes this responsible.
Whether 650 organizations counts as a sufficient vetting process for a model this capable is the question regulators and security researchers will debate.
Prompt injection robustness
Argon leads all frontier models on Gray Swan's Indirect Prompt Injection benchmark. This is the most practically relevant safety benchmark for enterprise deployment of autonomous agents.
Indirect prompt injection is the attack where malicious content in the agent's environment — a web page it reads, a document it processes, data returned by a tool — contains instructions that hijack the agent's behavior. It is how a web-browsing agent gets tricked into exfiltrating data, how a document processing agent gets redirected to different tasks, and how multi-step pipelines get subverted at the tool interface rather than the model interface.
Google's improvement comes from automated red teaming and adversarial training specifically targeting this failure mode. Argon also adds chain-of-thought and action monitors that halt execution when the model drifts outside user intent. Both of these are runtime controls in addition to training-time improvements — the defense is not just "train a more robust model" but "monitor what the model actually does."
For enterprise teams deploying agents with web access or document processing: this benchmark result is directly relevant to your threat model. Independently verify it on your specific workloads, but the fact that Google is specifically claiming and testing this capability reflects where the adversarial landscape is moving.
Pricing context
Introductory pricing: $2 input / $10 output per million tokens — matching GPT-6.1 Sol at DevDay yesterday. Standard pricing after the introductory period: $4 input / $20 output — matching Anthropic Opus 5.5's standard pricing. 95% cache discount matches Step 5 Preview's cache structure.
The introductory period has no announced end date. Watch for a pricing change announcement; the gap between $2 and $4 input (and $10 vs $20 output) is large enough to significantly affect cost models for sustained deployments.
Full pricing table with current context:
Model | Input ($/M) | Cached ($/M) | Output ($/M) | Notes Gemini 4 Argon (intro) | $2 | $0.10 | $10 | Intro period, no end date Gemini 4 Argon (standard) | $4 | ~$0.20 | $20 | After intro period GPT-6.1 Sol | $2 | $0.10 | $10 | GA Anthropic Opus 5.5 | $4 | $0.20 | $20 | GA GPT-6 Astra | $10 | $1 | $50 | GA StepFun Step 5 Preview | $1 | $0.05 | $2.70 | Preview, no license yet
At introductory pricing, Argon is price-equivalent to Sol with higher benchmark scores and 1M output capability. If the benchmark numbers hold under independent evaluation, this is the most favorable price-to-capability ratio in the frontier tier since September.
The Moonshot distillation connection
OpenAI disclosed this week that it banned Moonshot AI (makers of Kimi) on July 28 for running a large-scale adversarial distillation campaign — approximately 16,000 attempts to extract protected chain-of-thought reasoning from OpenAI's models. The ban came six weeks before the NSA/FBI/CISA advisory (September 8) that publicly named Moonshot as one of six Chinese AI companies running industrial-scale distillation campaigns.
The Moonshot case is relevant to Argon's launch because it documents the attack vector that Argon's prompt injection robustness is partly designed to resist. An adversarial distillation campaign does not only target the model through API extraction — it can also use prompt injection at scale to manipulate what the model returns, disguising the extraction as legitimate traffic. Argon's leading IPI benchmark position is Google's public response to this threat class.
When you can actually use it
Right now: if you are a Fairwind Program participant or Google internal team.
Soon: paid API customers and Google AI Ultra subscribers — no date given. Google has said "as soon as possible."
If you are planning to evaluate Argon for production use, the time to design your evaluation is now, before access opens. The DeepSWE benchmark at 77.9% and the Vals Index position suggest this is the strongest coding-and-reasoning model available at its price point — but all scores are vendor-reported, and the Long Decode Continuation architecture for 1M outputs needs testing in your specific workflow before you build dependencies on it.
The cybersecurity-first rollout is not just a safety theater exercise. Google is using the Fairwind cohort as a real-world evaluation environment for a model they are shipping without full guardrails. The feedback from 650+ organizations finding real vulnerabilities will shape the model's behavior before it reaches general availability. When Argon does open to developers, it will have been stress-tested in adversarial conditions at scale — which is more than most frontier models can say at general availability.



