Skip to content
Modelsstrong signalverified

Google announces Gemini 4 Argon for vetted cyber defenders only, at a $2 introductory input price

Google announced Gemini 4 Argon on September 30, 2026. The model is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google gives no date for developers, enterprises, or consumers. The introductory price is $2 per million input tokens and $10 per million output tokens, and it rises to $4 and $20 when the introductory period ends. The output limit grows to 1 million tokens from 64,000. The benchmark results in the announcement are Google's own reporting, and nobody outside the program can check them yet.

By Redakcija WebAiRadarPublished 4 min readwritten by a model
Image: Google

Gemini 4 Argon is Google's new frontier model, and for now you cannot call it. Google says access starts with trusted cyber defenders in its Fairwind Program and widens later, beginning with paid API customers and Google AI Ultra subscribers. What you can do today is plan around the published price and the new output limit.

Who gets access, and when

Google describes the release as phased. The first group is a set of trusted cyber defenders admitted through the Fairwind Program, which Google introduced earlier in September as a limited access program for governments and trusted partners. Google says it is taking part in the U.S. government's voluntary process for pre-release model access while it expands availability.

The announcement gives no date for anyone else. Google says it will make Argon available to developers, enterprises, and consumers as soon as it can, starting with paid API customers and Google AI Ultra subscribers. For trusted defenders and its own internal teams, Google says it will release the model without cyber guardrails.

What it costs

Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens are priced at 95% off the input price, which works out to $0.10 per million at the introductory rate. A footnote says the price after the introductory period is $4 per million input tokens and $20 per million output tokens. Google does not say how long the introductory period lasts.

The introductory numbers match a competitor's list price. OpenAI's API price list shows GPT-6.1 Sol at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens at the standard tier. This suggests Google set the introductory price to equal Sol's and kept the higher rate for later. If you build a long-term budget around Argon, the $4 and $20 figures are the ones to use.

The output limit grows to 1 million tokens

The output token limit rises to 1 million tokens, up from the previous limit of 64,000. Google's argument is that a model with room to generate hundreds of thousands of tokens in one run can solve a hard problem in a single pass. For you, the practical effect is on cost: a response that fills the limit costs $10 in output tokens at the introductory rate and $20 after it.

What Google reports on benchmarks

Every figure below comes from Google's announcement, and the model is not available for outside testing.

  • On DeepSWE v1.1, which measures long-horizon software engineering tasks, Google reports 77.9% and calls it a new state of the art.
  • On AutomationBench, Zapier's benchmark of end-to-end business tasks, Google reports 51.3% and first place.
  • On LVBench, which measures long video understanding, Google reports 91.7%.
  • On CWE-bench v1, which measures how well a model fixes security vulnerabilities, Google reports 68% and a tie for first place.
  • On the Vals Index, which weights finance, coding, legal, and tax work by share of U.S. GDP, Google says Argon leads but gives no score.

Cyber defense and internal use

Google says it trained Argon to find, validate, and patch critical software vulnerabilities on its own. According to the announcement, Wiz used the model in its Scan for Good initiative and found a critical vulnerability that exposed sensitive personal information in healthcare software used by hospitals. Google names neither the software nor its vendor.

The company also describes internal results. It says a team of Argon agents applied memory optimizations across Google's data centers that free more than 300 TiB of memory once rolled out. For libgav1, Google's open source video decoder, it says Argon agents replaced 32,000 lines of SIMD code in an existing Rust port. According to Google, the result runs 2.7 times faster than that port and produces identical video output.

Safeguards before a wider release

Google lists four areas it is still strengthening: misuse, prompt injection, misalignment, and the security of its own test environments. It says it monitors Argon's chain of thought and actions and stops execution when needed, and that it monitors the model's internal activations to spot misuse. Google also claims Argon leads Gray Swan's indirect prompt injection benchmark, again without a score.

„Safely releasing frontier capabilities at this level requires a phased approach.“
Koray Kavukcuoglu, Google DeepMind

Sources

BrandsGemini

Related

Screenshot of the OpenAI API dashboard, Project Settings, Text provenance tab: the Allow text watermarking switch is on, the Models menu reads All selected, and a Save button sits below.
Modelsstrong signal

OpenAI will watermark ChatGPT and Codex text in the EU, and API customers can opt in

OpenAI published its plan for text watermarking on October 5, 2026, in response to the EU AI Act. Over the coming weeks, eligible ChatGPT and Codex text output in the European Union will carry an invisible watermark called textGrain. In the API, watermarking is available worldwide from the same day for select models, and it stays off unless you turn it on. The detector is not public: OpenAI is limiting it to approved researchers and expert organizations.

OpenAIverified

LLM ROUTERSNone of 14router settings beat a random pick between two
Modelsmedium signal

A preprint finds six commercial LLM routers no better than a random pick between two models

A preprint submitted to arXiv on October 2, 2026, tested six commercial LLM routers in 14 settings. According to the authors, none of them beat a router that picks at random between Gemini 3.7 Flash and Opus 5 at the same cost, and one trailed it by 10.5 percentage points. The authors, who work at Fastino Labs, trace the gap to the way routers are evaluated and to rosters that hold too many models. The paper has not been peer reviewed.

arXiv, Fastino Labsverified

Microsoft's graphic for ThinkingBox: the name spelled in dotted letters on a yellow-green background covered with small blocks of horizontal lines.
Modelsmedium signal

Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row

Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.

Microsoft, Hugging Faceverified