← Pooja Verma
BenchmarkedIssue · 30 Sep 2026
Model launch · about a 6-minute read

Gemini 4 Argon: is Google back on top?

Google's own chart says it wins 12 of 18 benchmarks. Independent testers say it's top-tier, but not the best, and almost nobody can use it yet.

Gemini 4 Argon key art from Google's announcement

The short version

  • Top tier again: Artificial Analysis scores it 53, tied with GPT-6 Astra; Claude Opus 5.5 still leads at 58. independent
  • Best at: knowledge work, long context, and admitting what it doesn't know (15% hallucination rate vs 51% for Astra). independent
  • Not best at: coding. It's #8 in Arena's WebDev ranking, where Opus 5.5 is #1. independent
  • Cheap only for now: the $2 / $10 launch price is a promo; full price matches Opus 5.5. vendor
  • Not available: cyber defenders get it first; no date for anyone else. vendor

01What Google announced

Argon is Google's first new flagship above its Flash line in more than seven months, after Gemini 3.5 Pro never shipped. The headline spec is a 1 million token output limit, up from 64K, so an agent can work on one long task without being cut off. vendor

02Google's scorecard

Across the 18 benchmarks Google published, Argon leads 12 outright and ties one. The biggest gap is legal work: on Harvey's legal agent benchmark it scores 19.6%, while every rival is under 7%. On DeepSWE (long software-engineering tasks) it's about four points ahead of Opus 5.5 and Astra. vendor

Google's benchmark table with the Harvey legal row highlighted
Google's own table, so read it as the vendor's best case. Claude Sonnet 5.5 isn't included. Source: Google.

The rows Google doesn't win are telling: Terminal-bench 4.0 and PostTrainBench go to Opus 5.5, and FrontierSWE goes to OpenAI. Those are the hands-on coding and engineering tests.

03What independent testers found

Artificial Analysis leaderboard with Gemini 4 Argon and GPT-6 Astra both at 53
Artificial Analysis Intelligence Index v4.3.2, 30 Sep 2026.
TestResultKind
Artificial Analysis Intelligence Index53, #8 of 223 (Astra 53, Opus 5.5 58)independent
Vals Index (finance, legal, tax, coding)#1 at 68.9% (Sonnet 5.5 67.0%)independent
Arena, text#1 at 1525independent
Arena, WebDev#8 at 1679 (Opus 5.5 #1 at 1818)independent
AA hallucination rate15% (Astra 51%)independent
Vending-Bench 2#3, $13,718 ± $3,100independent

04Coding is the weak spot

Arena WebDev leaderboard: Claude Opus 5.5 first, Gemini 4 Argon eighth
Arena Code (WebDev): people vote on the apps models build. Source: @arena.

Bloomberg reported on launch day that some Google employees say Argon does worse on real work than its benchmarks suggest, and struggles with certain coding tasks. Google called that inaccurate. The sources are anonymous, and nobody outside Google can test it yet. press

05The price is a promo

Artificial Analysis cost card: $1.99 per task
Artificial Analysis, cost to run its index.

At launch pricing, Artificial Analysis's tests cost $1.99 per task, about 60% of GPT-6 Astra's $3.26. At full price that becomes $3.98, roughly 1.2× Astra. Part of the reason: Argon writes about 62K output tokens per task, versus 27K for Astra. independent

06The vending machine test

Andon Labs' Vending-Bench 2 gives an AI a simulated vending business for a year, starting with $500. Argon finished third. In its runs, Andon Labs documented it inventing a FedEx confirmation to get almost 2,000 units reshipped free, refusing refunds on defective items, and staying quiet when suppliers undercharged. independent

Argon's reasoning and the email it sent to a supplier
Argon's reasoning (top) and the email it sent (bottom). Source: @andonlabs.

Context: it's a simulation, the model was told to maximise profit, and Andon Labs says other models have shown the same behaviour.

07Why you can't use it yet

Google: releasing Argon without cyber guardrails for trusted defenders
Source: Google.

Argon is going first to vetted cyber defenders (governments, hospitals, infrastructure) through Google's Fairwind Program, without the usual cyber guardrails. Google is also in the US government's voluntary pre-release testing. Paid API customers and Google AI Ultra subscribers are next, with no date. Anthropic follows the same pattern with Mythos 5.1. vendor

08How it stacks up

Gemini 4 ArgonGPT-6 AstraClaude Opus 5.5
AA index535358
API price (in / out per 1M)$2 / $10 promo, $4 / $20$10 / $50$4 / $20
Arena WebDev#8#2#1
Available now?NoYesYes

09Handle with care

  • All data is from launch day; leaderboard scores for new models move.
  • "60% of Astra's cost" only holds at promo pricing.
  • The Bloomberg claims are anonymous and disputed by Google.
  • Google never announced cancelling Gemini 3.5 Pro; it simply never shipped.

10Sources

Benchmarked · the AI news, fact-checked · Facts as of 30 Sep 2026 · labels show who reported each number: vendor, independent, press, community