SPAWNSY

Gemini 4 Argon: Google Announces a Model Almost Nobody Can Use Yet. Price, Results and Access

Google announced Gemini 4 Argon, but only trusted defenders in the Fairwind program have access for now. The $2 and $10 introductory price doubles after launch, and every result comes from the producer.

AuthorTwenZySPAWNSY Editorial Desk
PublishedOctober 3, 2026
Read time5 min
SectionTech
Views360
Share
Gemini 4 Argon: Google Announces a Model Almost Nobody Can Use Yet. Price, Results and Access

Google announced Gemini 4 Argon, its new flagship model, on Wednesday, but the average developer cannot send it a single request yet. For now Argon goes to a narrow group of trusted cyber defence specialists through the Fairwind program, and a wider release is due "as soon as possible", first for paying API customers and Google AI Ultra subscribers. The price is known: $2 per million input tokens and $10 per million output tokens, but that is an introductory price, after which a price twice as high takes over. The launch has numbers and a price list, and no users who could check either.

Who can use Argon today

The post by Koray Kavukcuoglu of Google DeepMind describes a staged release. The first stage is trusted defenders in the Fairwind program, which we know from Gemini 3.8 Flash Cyber. For them and for its own teams Google is releasing Argon without cyber guardrails, so they can use the model's full ability to find, validate and patch critical vulnerabilities. At the same time the company is taking part in the US government's voluntary process for pre-release model access, and it is widening access gradually.

Google explains this as the need to strengthen safeguards before a broad release. It names four areas: protection against misuse in cyber and CBRN attacks, resistance to prompt injection through outside content (Argon is said to lead on Gray Swan's IPI test), monitoring for misalignment between what the user intends and what the model does by watching its reasoning and actions, and hardening the environments in which the model is tested. Google also mentions improving its techniques for monitoring the model's internal activations to catch attempted misuse.

A price with an expiry date

Argon will launch at an introductory price: $2 per million input tokens and $10 per million output tokens, with cached input tokens 95 percent cheaper, which is 10 cents per million. A footnote in Google's post says that after the introductory period $4 per million input and $20 per million output will apply. Google does not give an end date for the promotion.

The price matters most alongside the output limit. Argon has a 1 million token output limit, where earlier Gemini models had 64 thousand. A single response that uses the whole window costs $10 at the introductory price and $20 after it. For an agent working for hours on a code migration that is a reasonable cost. For someone sending hundreds of such tasks a day, the gap between the two prices decides whether it pays. OpenAI entered this tier earlier with GPT-6.1 Sol, which the company says approaches Astra on coding and computer use at one fifth of its price. Yahoo Finance reports that the rates are identical to Argon's. We wrote about OpenAI pulling another of its models, GPT-6.1 Astra, after safety testing in our piece on the cancelled GPT-6.1 Astra.

The results Google reports

Every number below comes from the producer's post and has not yet been independently verified. Google says Argon sets a new record on DeepSWE v1.1, a test of long software engineering tasks, at 77.9 percent, leads the Vals Index, which weights sectors by their contribution to US GDP, and ranks first on Zapier's AutomationBench with 51.3 percent.

One thing is missing from the post: competitors' results alongside. Google writes about state of the art and first places, but does not set Argon against specific OpenAI and Anthropic models, so the reader cannot tell by how much it is better. That matters because a few days ago, with Sonnet 5.5, we saw that in external measurement a new model's lead can be smaller than in the company's claims.

Google's own projects as evidence

More concrete than the results table are the internal uses, because they show the model at work. The quantum computing team used Argon to optimise the resources of subroutines, and the model beat the published baseline by 40 percent within minutes. Argon agents analysed profiling telemetry across the fleet and found memory optimisations that, once rolled out, will free more than 300 TiB, with an estimated 500 TiB to 1 PiB in total.

The third example is migrating code from C and C++ to Rust, from libraries of tens of thousands of lines such as re2 and libgav1 up to more than 800 thousand lines in the Fuchsia Zircon kernel. On the libgav1 video decoder, agents replaced 32 thousand lines of SIMD code in an existing Rust port through rounds of profile-guided experiments, and the result is a memory-safe decoder running 2.7 times faster than that port, with identical output. Google adds that the systems being rewritten go through automated and manual audits, emulation testing and review before rollout. These are producer claims about its own code, and nobody outside can check them.

Cyber defence: a model without brakes for the trusted

The cybersecurity passage sounds strongest. Wiz, a cloud security company, uses Argon in its Scan for Good program, which scans critical infrastructure for free. In an early test the model found a critical vulnerability in healthcare software used in hospitals worldwide, exposing personal data, which according to Google earlier frontier models had missed. On CWE-bench v1 Argon ties for first place with 68 percent.

That is an argument for the staged release and also its cost. A model that autonomously finds and patches vulnerabilities is as useful to an attacker as to a defender, which is why Google keeps the version without guardrails closed to a group of verified organisations.

Gemini 4 Argon is, today, an announcement with a price list rather than a product. The numbers in the post are impressive, above all the 77.9 percent on DeepSWE and the migration examples, but every one comes from the producer, and the model is available to a handful of cyber defence firms. In our view the reasonable reaction is to leave plans unchanged and wait for API access and for the first independent measurements.

The introductory price of $2 and $10 is a fine planning figure, but budgets should be built on $4 and $20, because those are the rates that will stay. With a 1 million token output limit, a single maximum task costs $20, and firms running hundreds of agents should work that out now, not after the promotion ends. Google has not given an end date for the introductory period, so the plan has to rest on the higher price.

On the plus side we put the openness: Google says plainly who gets the model, why and what comes next, without pretending that "soon" means "today". On the minus side, the lack of comparisons with competitors makes the results table hard to judge. We will write again when the model reaches the API and we can see how much of the 77.9 percent survives external tests.

Comments

Discussion

Join the conversation around this story.

0 entries

Join the discussion

Sign in to comment and reply to other readers.

Sign in

No comments yet

Start the discussion first.

Read next

All posts