Gemini 4's Real-World Performance Called Into Question by Google Employees
Google's Gemini 4 model has been making waves in the tech world with its impressive benchmark scores, but a recent report suggests that it may not always deliver in real-world use.
According to Bloomberg, some Google employees have found that Gemini 4's performance doesn't always line up with its benchmark scores. In particular, they've noticed that the model 'struggles to handle certain coding tasks'.
However, when questioned about these claims, a Google spokesperson maintained that it would be inaccurate to say that Gemini 4 is underperforming in areas such as coding. Koray Kavukcuoglu, head of Google DeepMind, previously stated that 'it's a certainty that we are always gonna be at the frontier,' suggesting that the model is still cutting-edge.
Gemini 4 Argon comes with some welcome upgrades, including an expanded output limit of up to 1 million tokens. It also boasts leading performance in cybersecurity defense and other use cases. The model will be priced at $4 per 1M input tokens and $20 per 1M output tokens, although introductory pricing will be half that.