On 30 September Google announced Gemini 4 Argon, its new frontier model (Google). This is what the published numbers say, what they do not, and how to be ready to test it on your own work.

Where it leads

In the benchmark table Google published with the launch (Google DeepMind), Argon is ahead on work most companies would recognise: knowledge work (Vals Index, 68.9% against 67.0% for Claude Opus 5.5), workflow automation (AutomationBench, 51.3% against 42.5%), finance agents (Vals Finance Agent v2, 65.4% against 58.9% for Claude Fable 5.1) and legal agents (Harvey’s benchmark, 19.6% against 6.7% or less). It also leads on long documents: on GraphWalks between 256,000 and 1 million tokens it scores 84.2%, against 71.8% for GPT-6 Astra. And it can write far longer answers, with an output limit of 1 million tokens, up from 64,000.

Where it trails

The same table shows Claude Opus 5.5 ahead on Terminal-bench 4.0 (66.4% against 57.4%) and PostTrainBench (49.3% against 45.3%), and GPT-6 Astra ahead on FrontierSWE v2 (65.5% against 55.0%), Terminal-Bench Science (68.1% against 57.6%) and OSWorld-2.0 computer use (72.6% against 69.2%). Google’s methodology notes that most Argon scores were computed by Google and most of the others are the providers’ own reported numbers. The independent check comes from Artificial Analysis, which places Argon at 53 on its Intelligence Index, level with GPT-6 Astra (Artificial Analysis).

Who can use it, and at what price

Not most teams, yet. Argon is rolling out first to cyber defenders in Google’s Fairwind Program, then to paid API customers and Google AI Ultra subscribers, with no date given. Google lists an introductory price of $2 per million input tokens and $10 per million output, with cached input 95% cheaper, rising later to $4 and $20. Artificial Analysis puts the cost of running its index at $1.99 per task for Argon during the discount, against $3.26 for GPT-6 Astra.

Our perspective: line up the test before the access

A leaderboard tells you where to look, not what to switch. The useful preparation is the same as with any new model. Pick twenty real tasks from the work where Argon claims its lead, such as analysing financial documents, reviewing contracts or answering from very long files. Write down what a good answer looks like, and run your current model on them now. When Argon reaches the API, the same set gives you a comparison on your own work, with cost per task next to quality, in an afternoon. Choosing the model per job rather than per vendor is the habit that pays off as frontier models keep arriving.