Google Gemini 4 Argon Outperforms Top AI Models in Key Benchmarks
Google has announced Gemini 4 Argon, its long-awaited flagship model, which beats OpenAI's and Anthropic's top models in most benchmarks. The company is taking a phased approach to releasing the model, citing a need to gather feedback from early testers and iterate on its guardrails before making it available outside of the Fairwind Program.
According to Google's benchmarks, Argon takes top billing or ties with other models in 13 out of 18 tests. However, the results are mixed when it comes to coding abilities, with Argon scoring highly on some tasks but trailing behind others by up to 10 points.
Argon excels in knowledge work, particularly in areas such as automation and graph walks. It also shows promise in cybersecurity, autonomously finding, validating, and patching software vulnerabilities. Google is releasing the model without cyber guardrails for Fairwind participants and its internal teams.
The company plans to make Argon available to paid API customers and AI Ultra subscribers first, before releasing it to developers, enterprises, and consumers. The pricing will be $2 per million input tokens and $10 per million output tokens during an introductory period, rising to $4 and $20 afterward.