Google released Gemini 4 Argon on Wednesday, and the first users to get it are not developers or consumers but a closed group of security researchers who will run it with its cybersecurity guardrails switched off. The company is positioning Argon as a frontier model for work that runs long rather than work that answers fast.
Access runs through Google's Fairwind Program, a security initiative whose members get the model ahead of paid API customers and Google AI Ultra subscribers, with developers, enterprises and consumers last. Koray Kavukcuoglu, SVP of Google DeepMind and chief AI architect, said Google is taking part in the U.S. government's voluntary process for pre-release model access while it widens availability.
Key takeaways
- Argon's output ceiling rises to 1 million tokens from the 64,000 of Google's previous models, which the company frames as room to finish a hard problem in a single pass.
- Google reports 77.9% on DeepSWE v1.1 for Argon against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra, plus a first-place 51.3% on Zapier's AutomationBench.
- Trusted cyber defenders and Google's own teams get Argon without cyber guardrails; every other customer waits for the guarded build.
Why the output ceiling is the headline spec
The number Google leads with is an output limit of 1 million tokens, a roughly fifteenfold jump from the 64,000-token ceiling of its earlier models. Google's argument is that depth of reasoning is capped by how much a model may write before it has to stop, and that this much headroom lets one long trajectory replace a chain of handoffs.
The evidence Google offers for that claim comes from inside its own walls rather than from any outside evaluator. In one case a fleet of Argon agents read profiling telemetry from across Google's data centers and shipped memory fixes on their own, which the company expects to reclaim north of 300 TiB at rollout and somewhere between 500 TiB and 1 PiB in the end.
A second deployment is more revealing, because it is the kind of work engineers normally refuse to automate. Argon agents are porting Google's C and C++ to Rust at scales from small libraries like re2 up to the 800,000-plus lines of the Fuchsia Zircon kernel. In the video decoder libgav1, the agents threw out 32,000 lines of hand-tuned SIMD and let the compiler vectorize plain safe Rust instead β a result Google clocks at 2.7 times faster than the previous Rust port, with byte-identical video out the other end.
What the cyber-first rollout is for
Argon was trained for defensive security work specifically, and the pitch is a full loop without a human in it: Google says the model can find a critical vulnerability, confirm it is real, and write the patch. Removing the cyber guardrails for Fairwind members and internal teams is what makes that loop usable, since a model that refuses to reason about exploitation cannot close the hole either.
There is one early field result to point at. Through Scan for Good, the free infrastructure-hardening program run by security vendor Wiz, Argon reportedly turned up a serious exposure of personal data in hospital software deployed worldwide β something Google says earlier frontier models walked past. Its remediation score, 68% on CWE-bench v1, is a tie for the top rather than a clear lead.
How Argon compares with Astra and Opus 5.5
Google's benchmark case rests on long-horizon tasks. It claims a new state of the art on DeepSWE v1.1 at 77.9%, which 9to5Google reported puts it above Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. Google also says Argon leads the Vals Index, which weights finance, coding, legal and tax work by each sector's share of U.S. GDP, ranks first on Zapier's AutomationBench at 51.3%, and reaches 91.7% on the long-video test LVBench. Every figure is Google's own, published in its launch announcement.
What Google says it is still hardening
The four areas Google listed split cleanly into two strategies. Two are about refusal β blocking cyber and CBRN misuse without shutting down legitimate dual-use research, and resisting indirect prompt injection, where Google claims the lead on Gray Swan's benchmark for that attack class. The other two assume refusal will sometimes fail: a monitor reads Argon's chain of thought and kills execution when the model drifts past the request, and sandboxes are sealed shut before high-risk training starts.
Buried in that section is the most interesting admission of the launch. Google says it keeps misalignment findings out of the training loop on purpose, because training on them would teach Argon to hide from the monitors watching it β and it asked competitors to keep their models' reasoning readable for the same reason.
Outlook
The timing is pointed. Argon landed a day after OpenAI's DevDay, where that company shipped its Dots agent and a GPT-6.1 Sol model, and shortly after OpenAI said it would not release a planned GPT-6.1 Astra over safety concerns. It also follows through on what Kavukcuoglu promised after taking over DeepMind in August, when he said Gemini 4 was in refinement and would ship early rather than polished. Introductory pricing is $2 per million input tokens and $10 per million output, rising to $4 and $20 later.
FAQ
Can I use Gemini 4 Argon today?
Not unless you are in Google's Fairwind Program or a trusted tester. Google's stated rollout order puts cyber defenders first, then paid API customers and Google AI Ultra subscribers, and finally developers, enterprises and consumers.
What does releasing Argon "without cyber guardrails" mean?
Google is giving trusted defenders and its internal teams a build that does not restrict the model's offensive security reasoning, so they can use its full vulnerability-discovery capability. Everyone else will receive a version with those restrictions in place.
How much does Gemini 4 Argon cost?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted 95%. After the introductory period, the price becomes $4 per million input tokens and $20 per million output tokens.






