Google has launched Gemini 4 Argon, a new frontier model the company says is built for complex, long-duration workflows across software engineering, legal and financial work, enterprise knowledge tasks, and cyber defense.

Argon is not being opened to the public all at once. Google said the model will follow a phased rollout, with the first access going to trusted cyber defenders through the Fairwind Program.
Pricing and initial rollout
Google set Argon API pricing at $2 per million input tokens and $10 per million output tokens. Cached input tokens are priced 95% lower than standard input tokens.
After early testing, Google said Argon will be opened in stages to developers, enterprises, and consumers. The first public users will include paid API customers and subscribers to Google AI Ultra.

Output cap jumps to 1 million tokens
The most closely watched change in Gemini 4 Argon is the increase in maximum output length, from 64,000 tokens to 1 million tokens in a single response.
Google drew a distinction between that 1 million-token figure and the context window. The company said the larger output budget allows the model to keep reasoning across a single task trajectory and generate results that run into the hundreds of thousands of tokens, or close to 1 million. It framed that capability around large code migrations, deep research, and multi-step enterprise workflows.
Some online commenters noted that many frontier models have output limits of about 128,000 tokens. By comparison, Argon’s 1 million-token single-generation ceiling stands out.
Benchmark results across software, finance, legal work, and video understanding
Google said Argon scored 77.9% on the DeepSWE v1.1 software engineering benchmark, which the company described as measuring a model’s ability to handle real, long-horizon software engineering tasks. Google said that score is state of the art.

In enterprise knowledge work, Google said Argon ranked first on the Vals Index. The index covers finance, programming, legal work, and tax, with weighting based on each sector’s contribution to U.S. gross domestic product.
Google also said Argon led on Vals Finance Agent v2, Harvey Legal Agent Benchmark, and Zapier’s AutomationBench. On AutomationBench, the model scored 51.3%.
Visual understanding was another area Google highlighted. The company said Argon can analyze professional charts, understand long videos, and act on information from multiple documents. On the long-video understanding benchmark LVBench, it scored 91.7%.
Already in Google’s internal engineering workflow
Google said thousands of employees are already using Argon in day-to-day work, including code debugging, algorithm design, and large codebase migration.

In quantum computing, Google said Argon helped researchers optimize space-time resources in quantum algorithms, defined as the product of qubit count and gate operation count. In one test, the company said, Argon improved a published baseline result by 40% in a matter of minutes.
In a data-center memory optimization project, an Argon agent analyzed performance monitoring data across Google’s network and automatically identified and implemented optimization plans. Google said those changes have already freed more than 300 TiB of memory, with expected total savings of 500 TiB to 1 PiB.
Argon has also been used in Google’s internal migration of C/C++ codebases to Rust, including core libraries such as re2 and libgav1, and extending to the Fuchsia Zircon kernel, which spans more than 800,000 lines of code. Because the work touches critical infrastructure, Google said the migrations still go through automated audits, human review, and simulation testing.
In the open-source video decoder libgav1, Google said an Argon agent rewrote about 32,000 lines of SIMD code through multiple rounds of performance testing and compiler analysis. The new Rust version, the company said, runs 2.7 times faster than the original Rust version while keeping video output unchanged.
Cybersecurity is one of the first deployment targets
Cybersecurity is one of the first areas where Argon is being put to work. Google said the model can independently discover, verify, and fix critical software vulnerabilities.
Cybersecurity firm Wiz has already used Argon through the Scan for Good program to check public infrastructure for high-risk exposures at no cost. Google said Argon identified a serious vulnerability affecting medical software used by hospitals around the world, one that could have exposed sensitive personal information. According to Google, earlier frontier models did not detect that risk.
On the CWE-bench v1 vulnerability remediation benchmark, Argon tied for first place with a score of 68%.

Google also said that in internal vulnerability testing across 20 programming languages, as well as in black-box penetration testing without source code, Argon outperformed Gemini 3.8 Flash Cyber.
Safety controls and questions from real-world use
As model capabilities increase, Google said it is tightening safety controls at the same time. Argon restricts requests tied to cyberattacks and to chemical, biological, radiological, and nuclear misuse. The company said it uses internal activation monitoring, automated red-team testing, and human evaluation to identify risks.
For indirect prompt injection attacks, Google said it improved the model’s defenses through adversarial training and automated testing. The company also put in place monitoring for model misalignment, tracking reasoning processes and action traces, stopping task execution when necessary, and isolating high-risk training and evaluation environments.
Still, benchmark performance and real work do not always line up. Bloomberg reported that Google internally still has doubts about Gemini 4’s performance on key tasks such as coding. People familiar with the matter said that while the model performs strongly on industry benchmarks, some coding tasks remain difficult to complete reliably in actual employee use.

Whether Argon can deliver on long-horizon reasoning and complex task execution will need more testing in real settings.
Part of a U.S. government pre-release process
Google said it is participating in a voluntary pre-release model process promoted by the U.S. government. The first testers will help the company evaluate the model and refine its safety protections.
Once early testing is complete, Argon will be opened step by step to developers, enterprises, and consumers, with paid API customers and Google AI Ultra subscribers among the first public users.

