Google DeepMind has launched Gemini 4 Argon, a new flagship model built for long-horizon, complex tasks spanning software engineering, finance, legal work, and cybersecurity.
Initial rollout is limited to trusted security teams
The model is not available to the general public at this stage. Google said the first wave of access will go only to trusted cybersecurity teams through the Fairwind Program. It plans to expand availability later to paid API customers and Google AI Ultra users.
Output cap jumps from 64K to 1 million tokens
One of the biggest changes in Argon is response length. The maximum output for a single response has increased from 64K tokens to 1 million tokens, giving the model room to sustain reasoning and carry out longer sequences of steps within one task.
Google said it has already used Argon internally for large-scale code migration, algorithm optimization, and data center memory optimization.
Benchmark results show mixed but strong performance
In Google’s official evaluations, Argon posted 77.9% on DeepSWE v1.1, beating GPT-6 Astra at 74.1% and Claude Opus 5.5 at 74.2%. It also ranked first on Vals Index and AutomationBench.
Argon did not lead across every benchmark. On FrontierSWE, Terminal-Bench 4.0, and some scientific tasks, it still trailed GPT-6 Astra or Claude Opus 5.5.
Cybersecurity focus and API pricing
Google placed heavy emphasis on Argon’s cybersecurity capabilities. The model can automatically find, verify, and patch software vulnerabilities. Versions delivered to trusted defensive teams may even have some cybersecurity restrictions removed to unlock fuller capability.
Before the formal rollout, Google said it is still testing for abuse, prompt injection, and agent overreach risks.
At launch, API pricing is set at $2 per million input tokens and $10 per million output tokens. Google said those rates will later double to $4 and $20.

