Gemini 4 Argon Review: 1M Output Tokens, a 15% Hallucination Rate, and a Controlled Release (2026)
Quick verdict
Gemini 4 Argon, Google's first Gemini 4 model, launched October 1 at an introductory $2/$10 per million tokens with a 1M-output-token ceiling — eight times the typical frontier limit — and a 95% cache-input discount. Artificial Analysis measures it tying GPT-6 Astra on the Intelligence Index at 60% of the per-task cost, with a 15% hallucination rate against 54% for Astra. It reaches general users only through a staged release: first Fairwind Program cyber defenders, then paid API customers and AI Ultra subscribers.
What Shipped
Google announced Gemini 4 Argon on October 1, 2026 — the first model in the Gemini 4 series, roughly 333 days after the previous flagship, and one day after DeepMind's new lead confirmed Gemini 4 had entered post-training. It is aimed at three things: real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense.
The headline specification is not the context window. It is the output ceiling: 1M tokens per request, up from 64K. Most frontier models cap a single response at around 128K. Holding the distinction matters — context is one number, generation length is another, and Argon's is roughly eight times the norm.
Pricing, With an Expiry Date
| Rate (per 1M tokens) | |
|---|---|
| Input | $2.00 |
| Output | $10.00 |
| Cached input | 95% discount |
That is aggressive — below Claude Opus 5.5's $4/$20 and far below GPT-6 Astra's $10/$50. The catch is stated by Google itself: this is introductory pricing and rates will rise. Any cost model built today should carry an explicit assumption about when $2/$10 stops being true.
The Numbers That Matter
Artificial Analysis measured two results that separate Argon from the field:
- Intelligence Index parity with GPT-6 Astra at ~60% of the per-task cost. Not a capability jump — a cost-per-result move, which is where competition has actually been happening all quarter.
- A 15% hallucination rate, against 54% for both GPT-6 Astra and GPT-6.1 Sol. That is the lowest recorded among current frontier models, and it is a more interesting claim than any accuracy benchmark. For enterprise work in law and finance, factual reliability is the binding constraint — a 3.6x reduction in fabrication rate changes what work can be delegated.
Google's own benchmark table, which should be read as vendor-reported:
| Benchmark | Argon | What it measures |
|---|---|---|
| DeepSWE v1.1 | 77.9% (SOTA) | Real-world, long-horizon software engineering |
| Vals Index | 1st | Enterprise knowledge work, GDP-weighted across finance, coding, legal, tax |
| LVBench | 91.7% | Long-video understanding |
| AutomationBench | 51.3% | End-to-end business workflows across 47 tools |
| CWE-bench v1 | 68% (tied 1st) | Security vulnerability remediation |
Thirteen of Google's nineteen published benchmarks show Argon leading. The internal-use evidence is more concrete than the table: Argon has been used on quantum algorithm subroutines (reportedly beating a published baseline by 40% in minutes), on data center memory optimization (freeing over 300 TiB), and on C/C++ to Rust migrations — in the libgav1 case replacing 32,000 lines of SIMD code with a memory-safe decoder that ran 2.7x faster than the original Rust port.
On security, Google says Argon autonomously found, verified, and patched critical vulnerabilities — including a sensitive-data exposure in medical software used by hospitals worldwide that earlier frontier models had missed.
The Release Is Staged, and That Is the Story
Argon does not launch to the public. It goes first to vetted cybersecurity defenders through the Fairwind Program, participated in a US government pre-release safety evaluation, and will then reach paid API customers and Google AI Ultra subscribers. General availability has no announced date.
Read that alongside the timing. Argon shipped during the most turbulent month in frontier-model security history — a state subpoena, a Senate summons, and a fresh set of misalignment disclosures. A model whose defining capability is finding exploitable vulnerabilities in real software arriving with a gated, verified-user release is a defensible posture on the merits — and also a positioning play: Google gets to be the vendor that shipped the security model quietly, to defenders, after government review.
The trade-off for buyers is straightforward. You cannot adopt what you cannot access. Teams outside the Fairwind Program and not on AI Ultra have a vendor benchmark table, not a model.
What To Watch
- General availability and the pricing step-up — both are unannounced, and both determine whether Argon is competitive or niche
- Independent hallucination replication — 15% is extraordinary if it holds; AA's measurement is the only non-vendor number so far
- The rest of the Gemini 4 series — Argon is the first of a generation, and Google has historically followed flagships with cheaper tiers (see Nano Banana 2.1 for the same pattern on the image side)
- Whether the 1M output ceiling becomes a category requirement — if long single-trajectory generation proves useful in practice, Anthropic and OpenAI will be pushed to match it
Summary
Gemini 4 Argon is the most interesting frontier release of the quarter on two dimensions: a 1M-token output ceiling that changes what "one task" means, and a 15% hallucination rate that would matter more than any capability benchmark if it replicates independently. The $2/$10 introductory pricing makes it the cost leader among flagships — briefly, by Google's own admission.
The complication is access. This is a model launched to defenders rather than developers, and until general availability lands, the honest assessment is: the vendor evidence is strong and unusually specific, the independent evidence covers exactly two numbers, and most teams cannot test it yet.
Related Articles
Frequently Asked Questions
What is Gemini 4 Argon?+
What does the 1M output token limit actually mean?+
How much does Gemini 4 Argon cost?+
How does it compare to GPT-6 and Claude?+
Why is the release restricted?+
Pros
- 1M output tokens per request — roughly eight times the 128K ceiling most frontier models impose
- Introductory $2/$10 per million tokens with a 95% cache-input discount, well below comparable frontier pricing
- Artificial Analysis measures per-task cost at about 60% of GPT-6 Astra while tying it on the Intelligence Index
- A 15% hallucination rate, versus 54% measured for GPT-6 Astra and GPT-6.1 Sol
- DeepSWE v1.1 at 77.9% and AutomationBench at 51.3% are state of the art, not incremental
Cons
- Staged release: Fairwind Program partners first, then paid API customers and AI Ultra subscribers; general availability is unannounced
- Introductory pricing is explicitly temporary — Google says rates will rise
- Vendor benchmarks dominate the evidence base; independent replication is thin this early
- Safety posture is conservative by design, which narrows what the model will do
- Cyber safeguards assume a vetted, defensive user; broad access is gated behind a verification program
Keep reading