All figures below come from Google's own announcement, so they are self-reported results rather than independent evaluations.
1. What Is Gemini 4 Argon?
Gemini 4 Argon is Google's newest frontier model, built to sustain deep reasoning across complex, long horizon workflows. Google says it delivers frontier performance in three areas: real world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.
A frontier model is one at or near the leading edge of what current AI systems can do. A long horizon workflow is a task with many dependent steps, such as migrating a large codebase or researching a financial question across many documents, where the model has to stay on track for an extended period.
Argon is not yet widely available. It is rolling out first to a set of trusted cyber defenders through Google's Fairwind Program. Google describes a phased approach and says it is taking part in the U.S. government's voluntary process for prerelease model access while it gradually expands availability.
2. Why Does a 1 Million Token Output Limit Matter?
2.1 What is an output token limit?
A token is a small chunk of text, often a word or part of a word. The output token limit is the maximum number of tokens a model can generate in a single output.
| Model | Output token limit |
|---|---|
| Previous Gemini limit | 64K tokens |
| Gemini 4 Argon | 1 million tokens |
Google says Gemini 4 Argon increases the output token limit from 64K to 1 million tokens, giving the model much more room to generate long outputs and sustain complex, multi-step workflows.
2.2 How does a larger output limit help in practice?
When a model has room to think at length, it can work through a hard problem in one continuous pass instead of being cut off partway. Google says this headroom lets the model generate hundreds of thousands of tokens in a single trajectory, which adds depth of reasoning for difficult problems.
For professionals, the practical value is in tasks that are naturally long, such as large code changes, detailed legal drafts or multistep research. A bigger output window does not guarantee better results on its own, but it removes a ceiling that previously forced these tasks to be split into smaller pieces.
3. How Is Google Using Gemini 4 Argon Internally?
Google reports that Argon is already powering internal workflows, with thousands of its employees pointing to strengths in specialized coding, deeper research and writing quality. The company shares three examples.
| Use case | What Argon did | Reported result |
|---|---|---|
| Quantum algorithm optimization | Optimized spacetime resources (qubits × gates) of key subroutines | Beat the published baseline by 40% in minutes |
| Memory efficiency | Agents analyzed fleet wide profiling data and applied optimizations | Over 300 TiB freed, 500 TiB to 1 PiB estimated in total |
| Codebase migration | Agents are migrating C and C++ code to Rust | Scale from tens of thousands to 800K+ lines |
3.1 What does each example involve?
In the quantum example, spacetime resources means qubits multiplied by gates, a common way to measure the cost of a quantum routine. The memory work used a team of agents. An agent, in this context, is an AI system that takes actions toward a goal rather than only answering questions.
The migration work targets libraries such as re2 and libgav1 and extends up to the Fuchsia Zircon kernel. Rust is often chosen for its memory safety guarantees. Because many of these systems are critical, Google says the rewrites are going through automated and manual auditing, emulation testing and review before reaching production.
3.2 What happened with libgav1?
Libgav1 is Google's open source software for decoding video. Agents took an existing Rust port and replaced 32K lines of SIMD code. SIMD refers to instructions that process multiple data points at once, which matters for video speed.
The agents ran many rounds of profile guided experiments, studied the compiler's output, and wrote safe Rust that the compiler could vectorize automatically. The result, according to Google, is a memory safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++ version.
4. How Does Gemini 4 Argon Perform on Benchmarks?
4.1 What do the reported results show?
| Benchmark | What it measures | Reported result |
|---|---|---|
| DeepSWE v1.1 | Real world, long horizon software engineering | 77.9%, new state of the art |
| Vals Index | Economic impact across finance, coding, legal and tax | Leading performance |
| Vals Finance Agent v2 | Multistep financial research | Leading performance |
| Harvey's Legal Agent Benchmark | Legal research and drafting | Leading performance |
| AutomationBench | End to end execution across core business functions | 51.3%, ranked first |
| LVBench | Long video understanding | 91.7%, state of the art |
The Vals Index weights each sector by its contribution to U.S. GDP. AutomationBench is Zapier's benchmark. Google also points to visual understanding as a strength in knowledge work, including professional chart analysis, identifying details from long videos, and taking action based on a series of documents.
4.2 How should you read these benchmarks?
A benchmark is a standardized test that lets different models be compared on the same tasks. It is a useful signal, but it has limits. Scores show performance on a defined set of problems, which may not match the shape of your own work. Results published by a model's developer also benefit from independent replication over time.
A reasonable approach is to treat benchmark numbers as a starting point, then test the model on a small sample of your own tasks once it becomes available.
5. What Can Gemini 4 Argon Do for Cybersecurity Defense?
5.1 What defensive capabilities does Google describe?
Google says it trained Argon to be highly capable at cybersecurity defense. According to the announcement, the model can autonomously find, validate and patch critical software vulnerabilities. For trusted defenders and Google's internal teams, Argon will be released without cyber guardrails so they can use its full defensive capabilities.
5.2 What early results has Google shared?
Wiz is using Argon through its Scan for Good initiative, a program that protects critical public infrastructure for free by finding and remediating high risk exposures. In an early demonstration, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide. Google says this was a severe risk that previous frontier models had missed.
| Benchmark | What it tests | Reported result |
|---|---|---|
| CWE-bench v1 | Remediation of security vulnerabilities | 68%, tied for first place |
| Google internal vulnerability benchmark | Finding exposures in complex codebases | Wide range found across 20 programming languages |
| Wiz black box penetration testing | Analyzing live web systems without source code | Outperforms 3.8 Flash Cyber |
On the Wiz benchmark, Google says Argon does better at discovering the attack surface, identifying vulnerabilities and producing proof of concept evidence to validate them.
6. What Safeguards Is Google Putting in Place Before Broad Release?
Google says it is strengthening safeguards across four areas before making Argon widely available.
| Area | Risk addressed | What Google says it is doing |
|---|---|---|
| Defending against misuse | Cyber and CBRN attacks | Refusal of harmful requests, activation monitoring, red team testing |
| Defending against prompt injection | Hijacked model behavior | Automated red teaming and adversarial training |
| Monitoring for misalignment | Actions beyond user intent | Monitoring chain of thought and actions, stopping execution when needed |
| Hardening systems | Insecure test environments | Isolating and sealing sandboxes before high risk work |
6.1 How does Google approach misuse?
The model is designed to refuse harmful requests tied to cyber attacks and chemical, biological, radiological and nuclear (CBRN) threats, while preserving legitimate dual use scientific research. Dual use research is work that has both beneficial and potentially harmful applications.
Google says it is improving techniques that monitor the model's internal activations to spot misuse, and that internal and external red teams tested the safeguards using manual and automated attack methods. A red team is a group that tries to break a system on purpose so weaknesses can be found and fixed.
6.2 What is indirect prompt injection?
It happens when malicious instructions hidden in content, such as a web page or document, are used to hijack a model's behavior. Google calls Argon its most resilient model yet against these attacks and says it leads on Gray Swan's Indirect Prompt Injection benchmark. It also notes that these attacks require constant vigilance and multiple layers of defense.
6.3 What is misalignment monitoring?
Misalignment, in this context, means a model pursuing a task in a way that goes beyond what the user intended. Google also says it used a similar system to monitor its training runs and alert an incident response team.
It took precautions against feeding those findings back into training, so that Argon's reasoning would not be shaped to evade monitoring. Google encourages the wider industry to preserve reasoning transparency so model thoughts remain useful for diagnosing misalignment.
6.4 Why does system hardening matter?
Safely testing powerful models requires secure environments. In line with its agent control roadmap, Google is isolating and sealing sandboxed environments before high risk training or evaluations begin. It also says it is committed to sharing these agent security practices with partners.
7. How Much Does Gemini 4 Argon Cost?
| Token type | Introductory price | Standard price |
|---|---|---|
| Input tokens | $2 per million | $4 per million |
| Output tokens | $10 per million | $20 per million |
| Cached input tokens | 95% off input price | Not stated |
Cached input tokens are inputs the system has already processed and stored for reuse, which is why they cost less.
8. When Is Gemini 4 Argon Available?
Argon is currently limited to trusted cyber defenders and early testers. Google says it plans to release to developers, enterprises and consumers as soon as possible, starting with paid API customers and Google AI Ultra subscribers. Google has not given a specific date.
9. Conclusion
Gemini 4 Argon is Google's attempt to build a model that can stay focused through long, demanding work, from large code migrations to financial research and vulnerability patching. The most notable details are the 1 million token output limit, the reported benchmark results, and the real world examples Google shares from its own engineering teams.
Equally important is the way Google is releasing it, with a phased rollout, government prerelease engagement and a focus on safeguards before broad access. For professionals, the practical next steps are to watch how access expands beyond trusted defenders, look for independent evaluations, and plan small tests on real tasks once the model becomes available.
