Google has unveiled Gemini 4 Argon, a new frontier model built to handle complex, multi-step workflows. The model is currently being deployed to select cybersecurity professionals through the company's Fairwind Program, with a broader release for developers and enterprises planned for a later date. Google stated it is participating in the U.S. government’s voluntary pre-release testing process to ensure safety before wider availability.
A primary feature of Gemini 4 Argon is its expanded output capacity, which reaches 1 million tokens. This represents a significant increase from the 64,000-token limit found in previous iterations. Google claims this headroom allows the model to perform deeper reasoning and solve intricate problems in a single trajectory, rather than requiring multiple fragmented prompts.
Internally, Google has utilized the model for large-scale engineering tasks, including migrating C/C++ codebases to Rust. In one instance involving the libgav1 video decoder, the model produced Rust code that reportedly runs 2.7 times faster than a previous port by enabling automatic vectorization. The model also assisted in optimizing memory usage across Google’s data centers, identifying potential savings estimated between 500 TiB and 1 PiB.
Gemini 4 Argon has achieved high scores on several industry benchmarks focused on professional domains. It recorded a 77.9% score on DeepSWE v1.1 for software engineering tasks and a 51.3% score on Zapier’s AutomationBench. Additionally, the model is reported to be the top performer on the Vals Index, which evaluates economic impact across finance, legal, and tax sectors, and it achieved a 91.7% score on the LVBench long-video understanding test.
Pricing for the model is set at $2 per million input tokens and $10 per million output tokens. Google is also offering a 95% discount on cached input tokens. The company continues to gather feedback from early testers to refine guardrails and performance before the model reaches a general audience.
