Quick Facts

  • Google unveiled Gemini 4 Argon on Sept. 30, 2026, its first flagship model since Gemini 3.1 Pro in February.
  • Google's internal benchmarks show Argon leading on 12 of 18 tests, but independent firm Artificial Analysis scores it level with GPT-6 Astra and behind Claude Opus 5.5.
  • Argon launches at $2.00 per million input tokens and $10.00 per million output tokens, with introductory pricing set to double after the initial period.

Google unveiled Gemini 4 Argon on Wednesday, claiming its new flagship AI model retakes benchmark leadership from OpenAI and Anthropic. The announcement ends a months-long gap in Google's frontier model lineup and follows the quiet cancellation of Gemini 3.5 Pro, which was expected in June but never shipped.

Across 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one. GPT-6 Astra leads on three and ties Argon on one. Claude Opus 5.5 leads on two. Google pitched the model at long-horizon work: software engineering, legal and financial research, and cybersecurity defense.

Google DeepMind's Koray Kavukcuoglu described Argon as "a model built to sustain deep reasoning across long-horizon workflows, with strengths in real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense." CEO Sundar Pichai said teams across Google already use it heavily, from coding to quantum computing.

Where Argon Leads and Where It Falls Short

On Google's own numbers, Argon scores 68.9% on the Vals Index for knowledge work, ahead of Claude Opus 5.5 at 67.0% and GPT-6 Astra at 63.1%. On the coding benchmark DeepSWE v1.1, Argon's 77.9% tops Opus 5.5 at 74.2% and GPT-6 Astra at 74.1%. On AutomationBench, Argon reaches 51.3% against Opus 5.5's 42.5%.

The picture shifts on tasks involving terminal agents and machine learning workflows. Claude Opus 5.5 leads Terminal-Bench 4.0 with 66.4% against Argon's 57.4%, and Anthropic also wins the PostTrainBench machine learning test. For enterprise buyers, model choice remains workload-dependent even with Argon in the mix.

Independent benchmarking firm Artificial Analysis reached a more cautious verdict. On its Intelligence Index, Gemini 4 Argon scores 53 points, placing it level with GPT-6 Astra but clearly behind Claude Opus 5.5, which scores 57.6 points. Anthropic's smaller Sonnet 5.5 model also outscores Argon at 56 points. Artificial Analysis also found Argon burns through more than twice as many tokens per task as GPT-6 Astra, which cuts into the apparent pricing advantage.

Internal Skepticism and a Stock Reaction

Bloomberg reported that some Google employees believe Argon performs well on industry-standard benchmarks but less well in actual use, particularly on certain coding tasks. Some employees believe Anthropic's and OpenAI's models are improving at a faster rate than Gemini, though others believe Gemini 4 has caught up with leading AI labs.

Google disputed that characterization, saying it would be inaccurate to describe Gemini 4 as underperforming in coding. Kavukcuoglu said he was encouraged by the model's performance. Alphabet shares pared earlier gains Wednesday following the Bloomberg report, closing up roughly 0.5% after trading up more than 2%.

Pricing and Availability

Argon launches at $2.00 per million input tokens and $10.00 per million output tokens through Google's API, exactly half the price of Claude Opus 5.5. Cached input tokens are discounted 95% to $0.10 per million during the introductory period. After the introduction ends, published prices rise to $4.00 input and $20.00 output per million tokens.

The rollout is limited. Pichai said the model is available to U.S. government users and a set of trusted cybersecurity organizations through Google's Fairwind Program. Broader availability has not been announced. The limited release follows a period of internal difficulty: Axios reported in July that poor morale at Google DeepMind contributed to delays that pushed Gemini 3.5 Pro out of its planned June launch window entirely.

Read more: Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.