NEWS
Google Sends Gemini 4 Argon to Cyber Defenders First
Google’s Gemini 4 Argon tops Vals AI’s main index, then withholds the model from developers so cyber defenders can patch what it finds.
Google on Sept. 30 released Gemini 4 Argon, a frontier model that leads Vals AI’s main index, only to selected cyber defenders. Paying developers do not get it yet.
Fairwind partners receive a copy with cyber guardrails stripped, so the model can hunt bugs at full strength. Broader access comes after that window, Google said, with no calendar date attached.
Fairwind Partners Get Argon Without Cyber Guardrails
Koray Kavukcuoglu, SVP of Google DeepMind and chief AI architect, said the company is rolling Argon to trusted cyber defenders through Fairwind while it stays in the U.S. government’s voluntary pre-release access process. Google launched that program on Sept. 2 around Gemini 3.8 Flash Cyber. Argon is now the model in that queue.
Fairwind already lists more than 650 participating partners, spanning governments, national cyber authorities, critical-infrastructure operators, Google Cloud customers, and security firms. Partners must keep the tools inside internal cybersecurity, incident-response, or penetration-testing teams and use multi-factor authentication.
Google trained Argon to find, validate, and patch vulnerabilities on its own, including black-box tests against live websites with no source code. For those trusted defenders and for Google’s own teams, the company said it is shipping Argon without cyber guardrails so they can use that full skill set.
FAIRWIND ACCESS RULES
- The cohort: More than 650 governments, infrastructure operators, cloud customers, and security partners are already in the program.
- Staff limits: Use stays inside internal cybersecurity, incident-response, or penetration-testing teams.
- The copy they get: Fairwind and Google teams receive Argon with cyber guardrails removed.
- The next gate: Developers, companies, and consumers wait on later paid API access and Google AI Ultra, with no public date.
That split is the launch. The same system that Google wants in hospitals and power networks is, in this first phase, a gated tool for people whose job is to break software before someone else does.
Introducing Gemini 4 Argon – our new frontier model.
It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program. pic.twitter.com/X8acOWJOSF
— Google DeepMind (@GoogleDeepMind) September 30, 2026
How Strong Is Gemini 4 Argon on Coding Tests?
Vals AI, which scores models on finance, coding, legal, and tax work weighted by each sector’s share of U.S. GDP, put Argon at 68.9%, No. 1 of 41 on the Vals Index. Claude Sonnet 5.5 is next at 67.04%. Each Vals Index run cost $15.68 for Argon, against $21.34 for Sonnet 5.5 and $32.14 for Claude Opus 5.5.
On Vals Finance Agent v2, Argon leads 73 models at 65.4%. Google’s own four-way card, which lines Argon up against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1, shows the same finance lead and a 77.9% mark on DeepSWE v1.1, a long-horizon software-engineering test. It also shows two coding boards Argon does not win.
GOOGLE’S FOUR-WAY SCORECARD
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% | 67.4% |
| CWE-bench v1 | 68.0% | 68.0% | 67.0% | 58.0% |
| Finance Agent v2 | 65.4% | 53.5% | 58.6% | 58.9% |
| Harvey Legal Agent | 19.6% | 5.4% | 3.8% | 6.7% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% | 56.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 66.4% | 57.9% |
CWE-bench v1, which scores vulnerability fixes, is a 68.0% tie with Astra. Harvey’s legal-agent test is the blowout: 19.6% for Argon against 5.4% for Astra and 3.8% for Opus 5.5. AutomationBench, Zapier’s end-to-end business-task score, is 51.3% for Argon, which Google called a first-place result.
A DeepSWE lead does not close the coding race. Astra still holds FrontierSWE v2 by 10.5 points, and Opus 5.5 holds Terminal-bench 4.0 by 9.0 points. Anyone treating Argon as a clean sweep on software is reading the press table, not the full card.
Argon Agents Are Already Inside Google’s Data Centers
Google is not waiting for a public API. Thousands of staff are already using Argon for debugging, research, and writing, Kavukcuoglu wrote, and some of the work is on production systems.
WHAT GOOGLE SAYS ARGON ALREADY DID
- Memory: Agent runs on fleet telemetry freed more than 300 TiB in Google data centers, with 500 TiB to 1 PiB estimated once the work is fully rolled out.
- Kernel port: Agents are moving C and C++ to Rust at scales from core libraries up to 800,000-plus lines in the Fuchsia Zircon kernel, under automated and human review before production.
- Video decoder: On libgav1, agents replaced 32,000 lines of SIMD in an existing Rust port and produced a build 2.7 times faster with identical video output.
- Quantum: On one subroutine, Argon beat a published baseline by 40% in minutes on spacetime cost.
That internal load is part of why a gated launch still has a customer. Google cannot rent this work from a rival without paying in both money and capacity, and Argon is already rewriting Google’s own stack while outsiders wait.
The spec people actually stop on is output length. Argon can emit 1 million tokens in one reply, up from 64,000, which lets a single run keep going on a long repair or a long audit instead of stitching a chain of shorter calls. Vals ran the model with a 1 million-token context window and a 262,144-token output cap, so independent scores do not yet reflect the full advertised ceiling.
A May Test Reached Three Real Companies
Google confirmed in mid-September that Gemini models, during a May capture-the-flag run by the security lab Irregular, reached three real companies after the test environment leaked onto the open internet. The made-up target shared a name with a live firm. In one run the model guessed passwords until it got in. In the other two it found credentials in public repositories and used them. Google said the model stopped once it saw it was on real systems, and that the firms were told.
Irregular notified Google in late July. Heather Adkins, Google’s vice president of security engineering, described the runs as a scoped test that spilled into the wild.
In a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.
Heather Adkins, vice president of security engineering, Google
OpenAI, Anthropic, and Meta disclosed related Irregular spills from the same testing bug. Google spoke after those labs, and after Irregular had already told the vendors. The May runs used older Gemini builds, not Argon. They still sit behind this launch: a lab that has watched its models walk onto live networks is now shipping a stronger cyber model to a closed defender list first.
FROM THE MAY SPILL TO THE ARGON GATE
- May 2026: Gemini models in an Irregular capture-the-flag test reach three real companies after unintended internet access.
- Late July 2026: Irregular notifies Google of the runs.
- Sept. 2, 2026: Google opens Fairwind with Gemini 3.8 Flash Cyber for vetted defenders.
- Sept. 18, 2026: Google confirms the May incidents and says the model stopped itself.
- Sept. 29, 2026: Sundar Pichai signs the White House Accord on Super Intelligence.
- Sept. 30, 2026: Gemini 4 Argon begins rolling out through Fairwind, without cyber guardrails for that cohort.
Wiz, already on Argon inside Fairwind, used the model in its Scan for Good work, which Google described as protecting critical public infrastructure for free. Google said Argon found a critical flaw that exposed personal data in healthcare software used by hospitals worldwide, a risk earlier frontier models had missed. Finding that bug is only a defense win if a fix is live before a wider drop puts the same skill in more hands.
Pichai Signed a Pledge Against Unintended Hacks
On Sept. 29, President Donald Trump hosted AI executives at the White House, including Pichai, and they signed a short voluntary text titled the White House Accord on Super Intelligence, also described as a Joint Commitment on Frontier Responsibilities. Signers included Pichai, Anthropic’s Dario Amodei, Meta’s Mark Zuckerberg, Nvidia’s Jensen Huang, OpenAI president Greg Brockman, and Elon Musk. Trump called the paper “morally binding.”
The first item asks companies to monitor models during training and deployment, including for cyber and biological risk, and to make sure those models do not hack or reach technical systems in unintended ways. The rest of the page points at internal teams, outside reviews, and board reporting. It names no penalties.
Google’s Argon post the next morning is, in practice, that pledge turned into a product schedule. The company said it is hardening sandboxes, watching chain-of-thought and actions so it can halt a run that drifts past the user’s intent, and tightening refusals on cyber and CBRN misuse under its Frontier Safety Framework. Tulsee Doshi, a senior director and head of product for Gemini, said, “We are seeing the guardrails be effective.”
Those layers are for a later, wider audience. The first humans on Argon are the ones Google will allow to use it as a pentest engine.
Rivals Still Lead Two Coding Benchmarks
Google had told investors it would ship Gemini 3.5 Pro by June. That model did not appear. Vals AI read the jump to Gemini 4 as a sign 3.5 Pro was not going to outrun the field, and Google’s own table still leaves FrontierSWE v2 with Astra at 65.5% against Argon’s 55.0%. Terminal-bench 4.0 stays with Opus 5.5 at 66.4% against 57.4%.
On a July investor call, Pichai had already conceded a lag and pointed at Gemini 4 as the reply. “There are many attributes at which we are still at the frontier,” he said then. Argon now leads a GDP-weighted work index and a long-horizon engineering test, and it still loses two coding boards that paying teams actually care about. The skipped 3.5 Pro looks, in that light, like a model that never earned a launch, not a delay that magically produced a clean win.
Broader Access Still Has No Public Date
Google said it will keep collecting Fairwind feedback and iterating on safeguards before Argon goes to developers, companies, and consumers, starting with paid API customers and Google AI Ultra subscribers. It did not say when. It is also in the government’s voluntary pre-release track, which can stretch a “soon” into a quarter.
Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access.
Koray Kavukcuoglu, SVP of Google DeepMind and chief AI architect, Gemini 4 Argon announcement
When the public API does open, the list price after the intro window is $4 per million input tokens and $20 per million output tokens. The intro rate is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off. That is cheap for a frontier model. It is also a price for a product most buyers cannot call yet.
Until Fairwind’s head start turns into patches on the software Argon is already reading, the lead on Vals is a private fact. Defenders have the unguarded copy. Everyone else has a scoreboard.
-
NEWS1 month agoGeneration Lab’s Secret Youth Shot Has a Copycat Problem
-
AUTO1 month agoTesla Raises Dual Motor Prices as Texas Builds Cybercabs
-
NEWS1 month agoOkta Stock Jumps 20% on McKinnon’s Identity Bet
-
LIFESTYLE2 months agoThe Nantucket Friendship Basket Boom Meets a Maker Shortage
-
BUSINESS1 month agoBristol Myers Quits Cellares as Autoimmune Doses Proceed
-
LIFESTYLE1 month agoLabor Day Mattress Sales Repeat a Familiar Holiday Discount
-
NEWS1 month agoGoogle Ends EU Spam Demotions but Keeps the Ranking Split
-
BUSINESS2 months agoPoland Closes Visa-Free Work for Three Fast-Growing Nationalities
