Gemini 4 Argon is Google's best model. You can't use it yet.
Google announced Gemini 4 Argon, its new frontier model, and only trusted cyber defenders can use it.
Google announced Gemini 4 Argon, its new frontier model, and only trusted cyber defenders can use it. Here is what Google's own 19-row benchmark table shows, what the independent leaderboard measured, and when developers might get it. Verdict: NEEDS REVIEW.
Read the written edition (English) ↗
What this video covers
- Gemini 4 Argon: a million output tokens, $2 in, and you can't use it
- Inside Google: Argon agents port C++ to Rust
- Google's own table: Argon tops 13 of 19 rows
- The cyber pitch: patches on its own, no guardrails for defenders
- The outside number: Artificial Analysis says 53, Opus 5.5 has 58
Transcript
Gemini 4 Argon: a million output tokens, $2 in, and you can't use it
0:00 You'd think a launch means you get the model. Google just launched Gemini four Argon, and unless you're a trusted cyber defender, you can't touch it. In this video: what does Google's table show? What did outside testers measure? And when do you get it? You've probably seen the hype videos already. I read the benchmark table instead, and two rows in it never made it into the
0:20 post's text. I'll get to them at the end. It's Thursday, October first, and this is The Daily Diff. Two stories today, and the big one is Gemini four Argon, which Google calls its next era of frontier intelligence. One answer can now run to a million tokens, up from sixty-four thousand, which is several novels in a single reply. The launch price is two dollars per million tokens in and ten out. That's exactly what OpenAI charges for GPT six point one Sol,
0:48 the model I covered yesterday, and Sol is one you can actually buy. Inside Google, it's already at work.
Inside Google: Argon agents port C++ to Rust
0:54 Google says Argon agents are porting C and C++ to Rust across the company, up to eight hundred thousand lines for the Fuchsia kernel. On Hacker News, one engineer remembered when Google's C++ team wouldn't even consider Rust and looked at Carbon instead. Now a model is doing the migration they argued about.
Google's own table: Argon tops 13 of 19 rows
1:12 Now the table. Google puts Argon next to GPT six Astra, Claude Fable and Claude Opus on nineteen rows, and Argon comes out on top in thirteen of them. The headline is DeepSWE, a long coding benchmark, at seventy-eight percent, about four points clear. The strangest win is legal drafting, where Argon scores twenty percent and the others stay in single digits. It's the best grade on a test that everyone fails, and on the slide-deck
1:38 benchmarks, which are undefeated, that counts.
The cyber pitch: patches on its own, no guardrails for defenders
1:41 The real pitch is security. Google says Argon can find, validate and patch critical vulnerabilities on its own, and that trusted defenders get it without cyber guardrails. Wiz used it to find a critical hole in hospital software that earlier models had missed. One small detail. Wiz belongs to Google, since a thirty-two billion dollar deal closed in March. So the black-box hacking test in the post is a Google company grading a Google model. And on the bug-fixing leaderboard, three models share sixty-eight
2:10 percent, and the bar on the far left belongs to Grok. So what did outside testers measure?
The outside number: Artificial Analysis says 53, Opus 5.5 has 58
2:15 The independent leaderboard I could find with Argon on it is Artificial Analysis. It scores Argon fifty-three on its intelligence index, on the high setting, which ties it with Astra and Fable. Claude Opus five point five sits five points ahead. A week ago I told you Opus took the top spot there, and it's still holding it. The fair part for Google is the bill. An Argon run costs about two dollars per task on that index, roughly a third of what Opus costs.
When do you get it: "So hold tight"
2:42 And when do you get it? Google won't give a date. The post promises access for developers, enterprises and consumers as soon as possible, with paid API customers first. Sundar Pichai's post on X says, so hold tight. Until then, access runs through Fairwind, the program Google started a month ago with Gemini three point eight Flash Cyber. It has over six hundred fifty partners, and they may only hand Argon to their
3:05 security teams and must track who uses it. Hacker News took it well. One commenter wrote, Gemini not beating the can't release a model allegations. Another predicted that models turn into vaporware, a bunch of numbers on a table. Meanwhile, Netlify rebuilt Edge Functions, which run about a billion
Netlify Edge Functions: V8 isolates to microVMs, about 5x faster
3:23 times a day. They moved from V8 isolates in a hosted service to Firecracker microVMs inside Netlify's own network, and a warm call dropped from up to forty milliseconds to about six. Hacker News pointed out that Cloudflare Workers are V8 isolates too and run much faster, so most of the win looks like the request no longer leaving the building. Another commenter worried about the snapshots, because cloned microVMs can share random number state, which is how you get two identical UUIDs.
3:50 A cold start, when a region has never seen your function, hits about one percent of calls and takes around nine milliseconds. And the microVM itself is Firecracker, which Amazon built for Lambda, so one commenter suggests you remember that the next time you curse AWS.
The two rows Google's post never mentions
4:06 Now, those two rows. On FrontierSWE and Terminal-bench, two coding tests the post's text never mentions, Argon comes last of four, on Google's own table. Terminal-bench puts an agent in a real shell, which is how most of us would use it, and Opus leads it by nine points. If you'd rather read this than hear me say it, the diff lands in your inbox every morning, free at the daily diff dot dev, link below.
Verdict: NEEDS REVIEW, the day I can call it
4:29 So, today's verdict on Gemini four Argon. NEEDS REVIEW. I'd stamp it that way because the independent number I found has it tied for second, and the two coding rows I care about have it last, so I'll review it properly the day I can call it. Subscribe, hit the bell, and tell me in the comments if you'd have stamped it differently. And that's the diff for today. I'm Niko from Axrisi. Merge responsibly.
Sources
- Googleblog.google
- Hacker Newsnews.ycombinator.com
- Fairwind Programdeepmind.google
- Artificial Analysisartificialanalysis.ai
- Sundar Pichaix.com
- Demis Hassabisx.com
- Google completes acquisition of Wizcloud.google.com
- Netlifywww.netlify.com
- Hacker Newsnews.ycombinator.com



