← All posts

Artificial intelligence

Google's First Frontier Model Since Gemini 3 Went to Cyber Defenders Before It Went to Customers. The Order Is the Announcement

Google says Gemini 4 Argon leads or ties on 13 of 18 benchmarks. Nobody outside a vetted programme can check, because the model went first to cyber defenders and a US government review process and only later to the people paying for tokens.

MAI
Google's official Gemini 4 Argon key art, the wide title graphic published with the company's own announcement post.

Google announced Gemini 4 Argon on September 30 — its first flagship model since Gemini 3 last November, and the first shipped under Koray Kavukcuoglu, who took over Google DeepMind after Demis Hassabis stepped aside in August. Google says Argon leads or ties on 13 of the 18 benchmarks it disclosed.

Almost nobody is in a position to check. Argon is not generally available. It went first to vetted cyber defenders through Google's Fairwind Program and into the US government's voluntary pre-release access process. Developers, enterprises and consumers get it "as soon as possible," starting with paid API customers and Google AI Ultra subscribers.

The order is the announcement. A company that has spent ten months being told it lost the frontier finally has a model it says takes the lead back, and it handed that model to incident responders and a government review process before it handed it to the people who pay for tokens.

What Google says it built

Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.

The headline engineering change is the output limit: one million tokens, up from 64,000. That is the difference between a model that drafts a file and one that can emit an entire migration in a single pass — Google says its own teams used Argon for large-scale codebase migrations, including C and C++ to Rust.

The disclosed comparisons, all measured by Google:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Opus 5.5
Harvey Legal Agent19.6%5.4%3.8%
DeepSWE v1.1 (software engineering)77.9%74.1%74.2%
Vals Finance Agent v265.4%53.5%58.6%
FrontierSWE v255.0%65.5%—
Terminal-Bench Science 0.157.6%68.1%—

On CWE-bench v1, a vulnerability-remediation benchmark, Google reports 68% and a tie for first. On Gray Swan's indirect prompt-injection test it reports a 0.7% attack success rate — the number that matters most for a model meant to be pointed at untrusted logs and tickets.

Defenders first is becoming the industry's release shape

This is not an isolated choice. On September 2 all three US frontier labs published on the same theme. Google launched Gemini 3.8 Flash Cyber through Fairwind, citing more than 650 partners including CrowdStrike, Datadog, Menlo Security, Palo Alto Networks and Snowflake. OpenAI said its Astra model had met the company's "Critical cybersecurity capability threshold" — meaning it can independently find and exploit previously unknown flaws — and routed access through a programme called Daybreak Blue. Anthropic kept Mythos 5.1 inside trusted-access programmes for cybersecurity and life sciences work.

What changed on September 30 is the scope. Gated access used to be the arrangement for a narrow security model. It is now the arrangement for the company's best general-purpose model. Google's supporting example is defensive: it says the security firm Wiz used Argon to uncover a critical vulnerability exposing personal information in healthcare software used by hospitals worldwide. That is a genuine argument for giving defenders a head start. It is also, read in the other direction, the reason the head start exists.

The four safeguards Google lists — misuse refusal, prompt-injection robustness, chain-of-thought monitoring for misalignment, and sandboxed execution — are described as things being strengthened before broad release, not things already finished. Google is saying, in the plainest terms it has used yet, that the general release is waiting on the safety work rather than on serving capacity.

The price is a promise about a model you cannot buy

Input per millionOutput per million
Introductory$2$10
Standard$4$20
Cached input (introductory)$0.10—

VentureBeat puts the introductory rate at roughly one-fifth of GPT-6 Astra's. That is aggressive, and it is also a price on a product with a waiting list. Until the API opens, it commits Google to a number without exposing it to a single customer's workload.

What to hold at arm's length

Every figure above is Google's own, produced on benchmarks Google selected, at a moment when no outside party can reproduce them. "Leads or ties on 13 of 18" is a count of a chosen set; the two results Google published where Astra wins are not close, and a model that trails by ten points on FrontierSWE v2 and Terminal-Bench Science 0.1 is winning on breadth, not on dominance. "As soon as possible" is not a date.

The useful thing to watch is not the next benchmark sheet. It is the length of the gap between the defenders' build and the public one. If it closes in weeks, the staged rollout was a launch sequence. If it runs for months, Google has established that a frontier model's first users are the people who secure systems and the government that reviews them — and the rest of the market gets the version that survives that review.

Sources: Google: Gemini 4 Argon — our next era of frontier intelligence, VentureBeat: Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic — but in limited release, Unite.AI: Google Announces Gemini 4 Argon Frontier Model for Coding, Cyber Defense, The Hacker News: Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs, Fortune: Demis Hassabis steps down from Google DeepMind CEO role amid a major AI leadership shake-up

Keep reading