
Targeted for Extreme Complexity
Argon moves beyond casual chat into end-to-end execution:
- Real-world software engineering & multi-file refactoring
- Enterprise knowledge workflows (legal contracts, financial risk modeling)
- Autonomous vulnerability patching and defensive cybersecurity
- Internal Google optimizations, from quantum computation tuning to Rust migrations
Google shared initial numbers placing Argon squarely back in the top tier against rival frontier models:
- AutomationBench (Zapier enterprise benchmark): #1 at 51.3%
- LVBench (Long-context video understanding): SOTA at 91.7%
- Reported 88.7 on DeepSWE v1.1 coding tasks
One of Argon's biggest architectural leaps is generation capacity:
- Context input remains massive (1M+ tokens supported)
- Output token ceiling leaps from the previous 64K up to a massive 1,000,000 output tokens in single continuous runs
- Ideal for generating complete codebases, auditing deep logs, or writing exhaustive reports end-to-end
Google is pricing Argon aggressively to undercut competing frontier options:
- Introductory: $2.00 / 1M input tokens | $10.00 / 1M output tokens
- Prompt Caching: Up to 95% discount on cached input tokens
- Post-Intro Standard: $4.00 / 1M input | $20.00 / 1M output
You won't find it broadly accessible in consumer apps just yet:
- Currently rolling out under early evaluation to vetted cybersecurity defenders via the Fairwind Program.
- Google noted participating in voluntary frontier safety review protocols before general public API endpoints go live.
After a cycle dominated by rapid updates from competitors, Gemini 4 Argon marks Google DeepMind's answer for developer and enterprise agents. Instead of raw conversational speed, Argon prioritizes deterministic tool use, persistent context coherence, and code correctness.
General availability in Google AI Studio and Vertex AI is slated to roll out soon following the closed red-teaming phase.