Google Announces Gemini 4 Argon Frontier Model With 1 Million Token Output
Google's own benchmark table is not an independent evaluation, and Argon trails competing models on coding and OS tasks like FrontierSWE v2 and OSWorld-2.0, so the performance lead is conditional on which benchmark is being read.
Key Facts
- Google announced Gemini 4 Argon on September 30, 2026.
- Gemini 4 Argon scored 68.9% on the Vals Index against GPT-6 Astra's 63.1% and Claude Opus 5.5's 67.0%.
- Gemini 4 Argon scored 84.2% on GraphWalks in the 256,000 to 1 million token input range.
- Gemini 4 Argon has an output cap of 1 million tokens, up from 64,000 tokens in the prior model.
- Gemini 4 Argon introductory pricing is $2 per 1 million input tokens and $10 per 1 million output tokens, rising to $4 and $20 after the introductory period.
Reporting from 3 sources: ASCII.jp, GameBusiness.jp, GIGAZINE.
Google announced Gemini 4 Argon on September 30, 2026, the first model in its Gemini 4 series. Google published benchmarks showing Argon leading on knowledge work, long-context processing, and business automation. On the Vals Index, Argon scored 68.9% against GPT-6 Astra's 63.1% and Claude Opus 5.5's 67.0%. On GraphWalks in the 256,000 to 1 million token range, Argon scored 84.2% against Astra's 71.8% and Opus 5.5's 66.8%. On AutomationBench, Argon scored 51.3% against Opus 5.5's 42.5% and Astra's 41.4%. Google did not release Argon to the public. It is being provided through the Fairwind Program to vetted cyber defense organizations and thousands of Google employees. Google plans to expand access in stages to paid API customers, Google AI Ultra users, developers, enterprises, and general users. Introductory pricing is $2 per 1 million input tokens and $10 per 1 million output tokens, rising to $4 and $20 after the introductory period.
Google CEO Sundar Pichai said on X that "with discussion building around the next model, we wanted to preview it as soon as possible." Several thousand Google employees already use Argon for coding, research, and writing.
- DeepSWE v1.1 (long-duration software development): Argon 77.9%, GPT-6 Astra 74.1%, Claude Opus 5.5 74.2%, Claude Fable 5.1 67.4%.
- Vals Finance Agent v2: Argon 65.4%, Fable 5.1 58.9%, Opus 5.5 58.6%, Astra 53.5%.
- Harvey Legal Agent Benchmark: Argon 19.6%, Astra 5.4%, Opus 5.5 3.8%, with all models still low in absolute terms.
- LABBench 2 (science): Argon 88.8%, Astra 85.4%, Opus 5.5 73.1%.
- Vibe Code Bench: 91.9%. LVBench (long video): 91.7%. GraphWalks under 128,000 tokens: 99.7%.
- FrontierSWE v2 (coding): Argon 55.0%, below Astra's 65.5% and Opus 5.5's 62.3%.
- Terminal Bench 4.0: Argon about 57%, below Opus 5.5's 60% and Claude Sonnet 5.5's 64%. Terminal-Bench Science 0.1: 57.6% against Astra's 68.1%. OSWorld-2.0: 69.2% against Astra's 72.6%.
- CWE-bench v1 (vulnerability fixing): Argon 68%, tied with Astra and one point above Opus 5.5's 67%. Argon ran in Antigravity, Astra in Codex, Opus 5.5 in Claude Code.
The output cap grows from 64,000 tokens to 1 million, about eight times the 128,000-token ceiling of Opus 5.5 and Astra. Google lists four large benchmark gaps against both rivals: AutomationBench, the Harvey legal benchmark, Vals Finance Agent v2, and GraphWalks. Artificial Analysis scored Argon 77.5% on AutomationBench-AA, about six points above Claude Sonnet 5.5's 71.3%.
Synthesized by Yomimono from the 3 cited sources below, including Japanese-language reporting where cited, then editorially reviewed before publishing.