Connect with us

NEWS

Google Holds Gemini 4 Argon for Vetted Cyber Defenders

Google launched Gemini 4 Argon on September 30 with 1 million output tokens, then limited it to Fairwind cyber partners with the guardrails off.

Published

on

Google opened Gemini 4 Argon on September 30, 2026, then kept it off the public Gemini dashboard and the paid API.

Fairwind cyber partners get it first. Their copy ships with the cyber guardrails removed, while everyone else waits for a date Google has not set.

Trusted Defenders Get Argon With the Guardrails Off

Koray Kavukcuoglu, SVP of Google DeepMind and Google’s chief AI architect, said a release at this level needs a phased rollout. The first phase is not a chatbot. It is a short list of trusted security teams, plus Google’s own staff, running a copy of Argon with the cyber brakes taken off.

DeepMind already works with over 650 partners globally in Fairwind, the program that now holds exclusive Argon access. The page names governments, national cyber agencies, hospitals, phone networks, energy firms, banks, and core software platforms as the groups it wants in the queue.

Fairwind is not a preview of the consumer Gemini app. Partners may run Argon on its own or inside CodeMender, DeepMind’s code-security agent, which finds flaws and writes patches so teams do not have to build their own test harnesses. Google says it will keep gathering tester feedback and tightening filters before developers, companies, and consumers see the model, starting with paid API customers and Google AI Ultra subscribers.

Safely releasing frontier capabilities at this level requires a phased approach.

Koray Kavukcuoglu, SVP, Google DeepMind, Gemini 4 Argon announcement

The public model is built to refuse harmful requests. The Fairwind copy is not. Kavukcuoglu wrote that trusted defenders and internal Google teams will get Argon without cyber guardrails so they can use its full defense skill. That is the launch. The scoreboard is marketing around it.

Dual-Use Work Only, and No Resale

DeepMind’s own Fairwind text states the bind in one line. Frontier models can find and fix critical flaws fast. In the wrong hands, the same skills become a threat. So the program tries to give high-priority defenders a head start before those skills leak into the open.

FAIRWIND ACCESS RULES

  • Permitted work: Partners may run authorized threat tests, reverse engineering, and malware analysis only for defense or academic research.
  • Banned work: Creating malware is not allowed, and dual-use jobs sit under review.
  • Who may log in: Only internal security, incident-response, or penetration-testing staff, with user-level sign-in and phishing-resistant MFA.
  • No resale: Partners may not share, redistribute, or sell access, and DeepMind runs background checks on applicants.

Groups that do not get in are told to use CodeMender with public Gemini models and other Google threat-defense tools. Academic labs that only want defensive tests may apply. Students are pointed to CodeMender on Google Cloud. When Argon is served as a managed model on Gemini Enterprise, DeepMind says it can run with zero data retention.

The Hospital Software Google Will Not Name

The proof Google chose for the gated launch is a bug in health software, not a public coding demo. Wiz, the cloud security firm named in the announcement, is already running Argon in its free Scan for Good initiative, which hunts high-risk exposures in public infrastructure and helps fix them at no charge.

Kavukcuoglu wrote that the model uncovered a critical flaw that exposed sensitive personal information across healthcare software used by hospitals worldwide, a risk earlier frontier models had missed. Google did not name the product. It did not say whether the hole is closed.

WHAT WE KNOW

  • The claim: Argon found a critical exposure of personal data in healthcare software used by hospitals around the world.
  • The method: Wiz used the model inside Scan for Good, which pairs AI scans with human researchers who confirm impact and notify the operator.
  • The skill: Google says Argon can find, confirm, and patch critical software flaws on its own.

WHAT IS UNCONFIRMED

  • The product: The announcement does not name the healthcare software.
  • The fix: Google has not said whether the issue has been patched.
  • Independent check: No outside lab has published a matching write-up of that worldwide hospital finding.

That unnamed hole is why the wait exists. A model that can walk a live site with no source code, map the attack surface, and spit out proof is a gift to a hospital defender and a gift to anyone who wants to copy the method. DeepMind’s cyber page says Argon can run black-box penetration of web systems from the public behavior of a site alone, then write validated patches so defenders can close holes before someone else finds them.

On CWE-bench v1, which scores the ability to fix security bugs, Argon ties for first at 68% with GPT-6 Astra. Claude Opus 5.5 sits at 67%. Google also says Argon beat Gemini 3.8 Flash Cyber, the earlier Fairwind model, on an internal test that spans codebases in 20 languages, and on Wiz’s internal live-web pentest that hides the source.

DeepSWE Hits 77.9%, FrontierSWE Does Not

Google is asking the rest of the market to judge a model it will not sell yet. The comparison it published is a selected 18-test card against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. Argon leads 12 of those tests outright, ties for first on 1, and trails on 5.

The headline coding win is DeepSWE v1.1, a long-horizon software-engineering test. Argon scores 77.9%, which Google calls a new state of the art. Opus 5.5 is at 74.2% and GPT-6 Astra at 74.1%, a 3.7-point gap. That is the job Google wants Argon known for: staying inside a real codebase for a long run instead of spitting a function and stopping.

Vals AI, which scores professional finance, coding, legal, and tax work, put Argon in first place on the Vals Index at 68.9%. It is Google’s first time on top of that board. Vals also has Argon first on Finance Agent v2 at 65.4%. Google’s own card adds AutomationBench at 51.3% and Harvey’s Legal Agent Benchmark at 19.6%, both first in that four-way set. Legal still fails four tasks in five, even at the top of the table.

GOOGLE’S FOUR-WAY SCOREBOARD

Benchmark Gemini 4 Argon GPT-6 Astra Claude Opus 5.5 Who leads
DeepSWE v1.1 77.9% 74.1% 74.2% Argon
Vals Index 68.9% 63.1% 67.0% Argon
AutomationBench 51.3% 41.4% 42.5% Argon
Harvey Legal Agent 19.6% 5.4% 3.8% Argon
FrontierSWE v2 55.0% 65.5% 62.3% GPT-6 Astra
Terminal-Bench 4.0 57.4% 58.2% 66.4% Claude Opus 5.5
CWE-bench v1 68% 68% 67% Tie, Argon and Astra

The coding crown splits as soon as the test changes. On FrontierSWE v2, Astra leads by 10.5 points. On Terminal-Bench 4.0, Opus 5.5 leads by 9.0 points. Anyone who lives in a terminal or an independent SWE board can read that as a model that wins Google’s long-run coding exam and loses the ones rivals already publish. Until Fairwind opens the door, nobody outside Google and its vetted list can rerun the jobs.

Long context is the other number Google can show without handing over an API key. Output is now capped at 1 million tokens, up from 64,000, so the model can keep thinking and writing in one pass. On GraphWalks items from 256,000 to 1 million tokens, Argon scores 84.2% against Astra’s 71.8%. On LVBench long video, Google reports 91.7%. Those figures travel well in a launch post. They do not let a shop migrate a kernel this week.

300 TiB Back and a Faster Video Decoder

The one customer already running Argon at scale is Google. Kavukcuoglu wrote that thousands of Googlers are using it for specialist coding, deeper research, and writing, and that agents are already in the company’s own stack.

WHAT ARGON ALREADY CHANGED INSIDE GOOGLE

  • Data center memory: Agents read fleet telemetry, applied memory fixes, and freed over 300 TiB once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
  • libgav1 decoder: Agents replaced 32,000 lines of SIMD code in a Rust port of Google’s open-source video decoder; the new memory-safe build runs 2.7x faster than that port, with the same video output.
  • Fuchsia kernel: C and C++ to Rust rewrites now range from tens of thousands of lines in libraries such as re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel, still under audit and emulation tests.
  • Quantum code: Researchers used Argon to cut spacetime cost (qubits times gates) on a bottleneck subroutine, beating a published baseline by 40% in minutes.

That internal ledger is the quiet half of the launch. Google is already spending Argon on memory, Rust, and quantum while the paid API is a footnote. The Fuchsia rewrite is not in production. Google says those large jobs still need automated checks, manual review, and emulation before they ship, which is the sane sentence in a post that otherwise reads like the model already recoded the company.

The memory number is the one finance teams will underline. Freeing over 300 TiB without new machines is a data-center win Google does not need Fairwind to collect. Paying customers are being asked to wait so a dual-use cyber skill can be boxed in. Google is not waiting on its own clusters.

Pichai Signed the Safety Accord on Tuesday

Sundar Pichai was at the White House on September 29, 2026, the day before Argon appeared. He signed the White House Accord on Super Intelligence, a one-page voluntary pledge also signed by President Donald Trump, Anthropic’s Dario Amodei, Meta’s Mark Zuckerberg, OpenAI president Greg Brockman, xAI’s Elon Musk, and Nvidia’s Jensen Huang.

Trump called the paper morally binding. The text asks labs for four layers: internal controls that watch training and deployment in areas that include cyber, biology, and chemical threats; an internal team to make those controls work; an independent outside auditor; and an independent board committee that reads the reports and sees that problems get fixed. The same week, Google said it is in the U.S. government’s voluntary process for pre-release model access.

THE WEEK THE MODEL WAS GATED

  1. September 29, 2026: Pichai signs the White House Accord on Super Intelligence with five other lab and chip leaders, pledging four layers of internal and outside review.
  2. September 30, 2026: Google announces Gemini 4 Argon, sends the unguardrailed cyber copy to Fairwind partners and internal teams, and holds the public API.
  3. Fairwind roster: DeepMind lists more than 650 partner groups already in the program, with Argon reserved for a set of them plus CodeMender.

Google is also hardening the copy the public will eventually get. The post lists four workstreams: misuse filters for cyber and for chemical and biological harm, including watches on the model’s internal activations; prompt-injection defense, where Google says Argon leads Gray Swan’s indirect-injection test; misalignment monitors that can halt a chain of thought mid-task; and sealed sandboxes for high-risk tests. Those are reasons to go slow. They are also reasons a hospital scanner with the brakes off cannot be a Tuesday download.

Gemini 4 Argon Has API Prices and No Ship Date

Google still printed a rate card. Argon will open at introductory $2 and $10 rates per million input and output tokens, with cached input at 95% off. After that intro window, the price becomes $4 per million input tokens and $20 per million output tokens. There is no public date attached to either tier.

The next groups named are paid API customers and Google AI Ultra subscribers. Developers, companies, and consumers come after the Fairwind tests. Kavukcuoglu’s post says “as soon as possible.” That phrase is doing a lot of work for a model Google is already using to free hundreds of tebibytes and rewrite video code.

The sticker is half the later rate, which will look cheap next to other frontier APIs once anyone can submit a prompt. Right now it is a number on a door that only a vetted security team can open. Shops that wanted Argon for legal research, finance agents, or a million-token coding pass are in the same line as hobbyists. The people at the front of that line are the ones Google believes should hold a bug-finding engine before the rest of the internet does.

Google built a model that hunts flaws in hospital software, patches C code, and writes for a million tokens at a stretch. It then gave the unguardrailed copy to Fairwind and to itself. The paid API still has a price, and it still has no day on the calendar.

Harry is the editor of SOMALI UPDATE, an independent title he owns and runs. Ten years in journalism, from reporter to editor, have settled into a set of verification habits he applies to every story. A quote is checked against the recording or transcript it came from. A statement attributed to an organisation is confirmed on that organisation's own channels before it is repeated. A figure is traced to the dataset or filing that first published it, and a photograph is checked for when and where it was actually taken. If any of those checks fails, the claim is left out or clearly marked as unconfirmed. Those habits cover the whole site, which reports news, business, technology, science and sports along with entertainment, lifestyle, travel, auto and gaming for readers around the world. Product claims in the technology, auto and gaming pages are tested in use where Harry can get his hands on the product. Corrections are published under a public policy and noted on the article. Readers who want to question a fact can write to support@somaliupdate.com.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending