On 30 September 2026, Google announced Gemini 4 Argon, its new frontier AI model, and is releasing it first to vetted cybersecurity defenders through its Fairwind Program rather than to the public. The model leads OpenAI’s GPT-6 Astra and Anthropic’s Claude Opus 5.5 and Claude Fable 5.1 on most of the benchmarks Google published, and it can write up to 1 million output tokens in a single response, up from 64,000 on previous Gemini models.
Google says Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens 95% cheaper. After the promotional period, the price rises to $4 and $20.
Wider access will start with paid API customers and Google AI Ultra subscribers, but Google has not given a date. It also has not announced India-specific pricing or availability. Meanwhile, the company says it is taking part in the U.S. government’s voluntary pre-release model access process.
The output limit is the main technical change. Frontier models now commonly accept 1 million tokens of input, but the amount they can write in one go, including their own reasoning, has been far smaller. Koray Kavukcuoglu, SVP of Google DeepMind, wrote in the announcement that this headroom “adds a new level of depth in reasoning.”
Independent benchmarker Artificial Analysis said it tested the full limit using a new Gemini API feature called Long Decode Continuation. That feature pauses a long response and resumes it across follow-up calls, so requests don’t time out.
Strongest at knowledge work, mixed at coding
Google’s benchmark table shows Argon’s clearest leads in professional tasks. It scores 68.9% on the Vals Index, which weights finance, legal, coding, and tax work by their share of U.S. GDP, against 67.0% for Claude Opus 5.5. On Zapier’s AutomationBench, which tests end-to-end business tasks, it scores 51.3% against 42.5% for Opus 5.5.
On Harvey’s Legal Agent Benchmark, the score is 19.6%, nearly three times the next-best score. Even so, that means it completes only about one task in five. Argon also achieves what Google calls a new state-of-the-art of 77.9% on DeepSWE v1.1, a test of long-running software engineering work.

The coding results are less consistent. Argon trails GPT-6 Astra by 10.5 points on FrontierSWE v2 (55.0% to 65.5%). On Terminal-bench 4.0, it comes last of the four models at 57.4%, nine points behind Opus 5.5. It also trails on PostTrainBench, a machine-learning engineering test, and on Terminal-Bench Science 0.1.
“Argon is a well-rounded model that has frontier capabilities across several domains,” Tulsee Doshi, head of Gemini products at Google DeepMind, told Axios. Artificial Analysis scored Argon at 53 on its Intelligence Index, matching GPT-6 Astra. It credited much of that to fewer hallucinations: Argon gave wrong answers on 15% of its AA-Omniscience questions, compared with 51% for GPT-6 Astra.
However, the firm also found Argon’s raw accuracy fell to 50%, five points below Gemini 3.1 Pro Preview. In other words, the model declines to answer more often than it knows more. Before the launch, Bloomberg reported that some within Google worried Argon would not match rival models from Anthropic and OpenAI.
Released to defenders without cyber guardrails
Google says it trained Gemini 4 Argon to find, validate, and patch critical software vulnerabilities on its own. Fairwind participants and Google’s internal teams get the model without the cyber guardrails that will apply elsewhere. Wiz, Google’s cloud security unit, is already using Argon in its Scan for Good programme, which finds risky exposures in critical public infrastructure for free.
According to Google, the model found a critical flaw exposing sensitive personal data in healthcare software used by hospitals worldwide, one that earlier frontier models had missed. Google did not name the software or say whether a fix has shipped.
On CWE-bench v1, which measures how well a model fixes security vulnerabilities, Argon ties for first at 68% with GPT-6 Astra and xAI’s Grok 4.7. Claude Opus 5.5 scores 67%, and Claude Fable 5.1 58%. Each lab’s model was tested inside its own coding agent: Antigravity for Gemini, Codex for GPT-6 Astra, and Claude Code for the Claude models. As a result, the leaderboard measures each model and its tooling together.

On two other security tests, Google compares Argon only with its own Gemini 3.8 Flash Cyber. Google’s charts show Argon scoring 85.8% on its internal vulnerability discovery benchmark, against 71.0% for Flash Cyber. On Wiz’s black-box penetration testing benchmark, it scores 70.9%, compared to 58.2%. The safety section of the announcement says Argon is designed to refuse requests to help with cyberattacks, while the same post confirms those refusals are switched off for the defenders receiving it first. Google says it monitors the model’s internal activations to spot misuse and its chain of thought to prevent actions that go beyond what the user intended.
Google has also been using Argon internally. The company says a team of Argon agents found and applied memory optimisations across its data centres, freeing more than 300 TiB of memory, with total savings estimated at 500 TiB to 1 PiB.
Argon agents are also migrating C and C++ code to Rust, including more than 800,000 lines of the Fuchsia Zircon kernel. Google says those rewrites are still being audited before they reach production. On libgav1, Google’s open-source video decoder, the new Rust version runs 2.7 times faster than an earlier Rust port.
Argon follows Google’s Gemini 3.8 Flash and 3.8 Flash Cyber release and effectively replaces the Gemini 3.5 Pro model Google announced at I/O in May. Google has not said when Argon will reach developers and consumers, which organisations are in the Fairwind Program, or when the introductory pricing ends. Artificial Analysis said the discount runs for at least one month.
Community Discussion
Join the conversation. Ask questions, share solutions, and help others.
Be the first to start the discussion!
