Announced September 30, 2026 · Rolling out in stages

Gemini 4 Argonlong research, written out in full

Google's first frontier model since Gemini 3, built for deep reasoning across long, multi-step work. It lifts the output ceiling from 64K tokens to 1M, and it leads Google's published numbers on knowledge work rather than on coding.

The Felo API is in closed beta — Gemini 4 Argon access is coming soon.

Gemini 4 Argon key art: the model name beside a large numeral 4 on a blue gradient
Google's announcement art. Argon ships under a new naming scheme rather than as an update to the Gemini 4.0 line, which is why the name carries no version number.
1M
Max output tokens
Up from 64K — a full report generated in a single trajectory.
$2 / $10
Introductory API price
Per million input and output tokens. Cached input is 95% off.
77.9%
DeepSWE v1.1
Google's reported score, ahead of Claude Opus 5.5 and GPT-6 Astra.
15%
Hallucination rate
The lowest measured among frontier models by Artificial Analysis.

Figures and charts on this page come from the organisations that published them — Google, Arena.ai, and Artificial Analysis — and are reproduced unmodified.

Rollout status

Argon is not generally available yet. Here is the order it arrives in.

Google is releasing Argon in stages, and the first stage is not the public one. Access inside Felo is being set up alongside this rollout.

Available now

Trusted cyber defenders

Selected partners in Google's Fairwind Program receive Argon first, and receive it without cyber guardrails.

Announced next

Paid API customers and Google AI Ultra

Google says wider access begins here. No date has been given, and a subscription alone does not grant it today.

Not yet scheduled

Developers, enterprises, and consumers

Follows the paid tiers. Google has not committed to a timeline for general availability.

Argon access inside Felo Search is in progress. This page will be updated when it opens.

What it is good at

A model built to finish long work, not to answer fast

Argon's published strengths cluster around reading, reasoning, and writing at length. That is a different shape from the coding-first frontier models it is being compared to.

Reasoning that survives a long task

Google positions Argon for multi-step work that runs for hours, where earlier models drift or stop early.

1M tokens of output in one pass

Google's stated reason for the increase: a model with room to think at length can solve a hard problem in one trajectory instead of many round trips.

Knowledge work leads, not coding

The widest reported gaps are on Harvey's Legal Agent Benchmark and Vals Finance Agent v2 — real legal and financial analysis, not synthetic tasks.

Long-context retrieval that holds up

On GraphWalks, which tests whether a model can actually use a 256K to 1M context, Google reports a lead of more than ten points over other flagships.

Long video understanding

LVBench, which measures comprehension over long video, is reported at 91.7% — the highest score Google lists for Argon.

Bar chart of Harvey's Legal Agent Benchmark: Gemini 4 Argon 19.6%, Claude Fable 5.1 6.7%, GPT-6 Astra 5.4%, Claude Opus 5.5 3.8%
Long-horizon legal work. Argon's 19.6% is roughly three times the next model's score, and it is the widest lead in Google's published set.
Bar chart of Vals Finance Agent v2: Gemini 4 Argon 65.4%, Claude Fable 5.1 58.9%, Claude Opus 5.5 58.6%, GPT-6 Astra 53.5%
Financial research and analysis, at 65.4% against a 53.5-58.9% field. A narrower lead than on legal work, but the same ordering.

Published numbers

Where Argon sits on Google's own benchmarks

These are vendor-reported results. Google has not released a thinking or effort setting for the headline scores, and no independent lab has replicated them yet.

Benchmark
What it measures
Argon
Nearest comparison
DeepSWE v1.1
Long-horizon software engineering
77.9%
Opus 5.5 at 74.2%, GPT-6 Astra at 74.1%
CWE-bench v1
Finding and patching vulnerabilities
68%
Ties for first place
AutomationBench
End-to-end business tasks, from Zapier
51.3%
Ranked first
LVBench
Understanding long video
91.7%
Highest score Google reports for Argon

Treat these as directional. Vendor benchmarks rarely survive contact with production workloads unchanged, and Bloomberg has reported that some Google employees are themselves sceptical that the benchmark lead transfers to real work.

Bar chart of DeepSWE v1.1: Gemini 4 Argon 77.9%, Claude Opus 5.5 74.2%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%
Long-horizon software engineering. The 77.9% headline figure, with both nearest competitors inside four points of it.
Bar chart of AutomationBench: Gemini 4 Argon 51.3%, Claude Opus 5.5 42.5%, GPT-6 Astra 41.4%, Claude Fable 5.1 31.4%
End-to-end business tasks, built by Zapier. Argon clears 50% where every other model in the comparison sits in the thirties or low forties.
CWE-bench v1 leaderboard: Gemini 4 Argon, Grok, and GPT-6 Astra tie at 68%, Claude Opus 5.5 at 67%, tapering to Inkling at 37%
Vulnerability discovery at Pass@1. Argon ties the top of a crowded board rather than opening a gap, which is why the table above calls it a first-place tie.
Google's full results table for Gemini 4 Argon across 16 benchmarks, grouped into knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding, and cybersecurity
The complete table from the announcement, with Argon's wins in blue and the cells where a competitor leads in grey. FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench, Terminal-Bench Science 0.1, and OSWorld-2.0 are the rows where Argon does not lead.
Bar chart of Artificial Analysis's AA-Omniscience hallucination rate across 36 models, sorted lowest first: Gemini 4 Argon (High) at 15%, then MiniMax-M3 at 18%, Kimi K2.5 at 26%, and rising to DeepSeek-V4 Flash at 96%
Artificial Analysis's independent measurement, and the one figure on this page that no vendor supplied: how often each model answers incorrectly instead of declining to answer. Argon's 15% is the lowest of the 36 models plotted; the scale runs up to 96%.
Arena.ai's Text Arena and Code Arena: WebDev leaderboards side by side: Gemini 4 Argon ranked 1st in Text Arena at 1,625, and 8th in Code Arena: WebDev at 1,679
The two rankings the paragraph above rests on, from Arena.ai's crowd-voted leaderboards. Argon tops Text Arena's 25 listed models; in Code Arena: WebDev it is eighth, behind four Claude models and three from OpenAI's GPT-6 line.

On Felo

What a research workflow looks like with a 1M-token writer

Felo pairs search with document understanding, so a long-output model changes what a single session can produce.

Research that ends in a document

Gather sources, then have the model write the synthesis in one pass instead of assembling it section by section across many prompts.

Whole document sets in one task

Long-context retrieval is the capability Argon is strongest at, which is what makes contract sets, filings, and long reports tractable.

Drafting where tone matters

Argon's reported profile — strong text, moderate coding — suits teams that need prose that reads as written, not assembled.

Cost

What Argon costs, and what changes later

Google has published an introductory rate and a higher list rate, without saying when the switch happens. Plan your budget on the higher number.

Token type
Introductory
Later
Input
$2.00 / M
$4.00 / M
Cached input
95% off
95% off
Output
$10.00 / M
$20.00 / M

For comparison, the introductory rate sits roughly level with GPT-6.1 Sol and Sonnet 5.5, and near a fifth of GPT-6 Astra. Those are list prices, not measured cost per task — Argon answers harder questions, so a single run can still cost more.

How to work with Argon on Felo

01

Bring the whole question

Give Felo the document set or the long thread, not a summary of it. Argon's advantage starts where the context stops fitting in a prompt.

02

Ask for the finished artifact

With a 1M-token output ceiling you can request the full report or the complete draft in one instruction rather than in sections.

03

Check it against the sources

Felo keeps citations attached to the answer, which is what makes a long model-generated document reviewable.

Argon access in Felo Search is in progress. Until it opens, every other model in the Felo lineup is available now.

Open Felo AI Search

Gemini 4 Argon: common questions

Not generally. Google released Argon first to selected partners in its Fairwind Program, and says wider access will start with paid API customers and Google AI Ultra subscribers. No date has been announced for that stage, and a paid subscription does not grant access today.

Try Google's frontier models on Felo

Search, read, and write with the model that fits the task. Argon joins the lineup when its rollout reaches Felo.

Open Felo AI Search

No credit card required to use Felo AI Search.