Google's Gemini 4 Argon Model Shows Mixed Results in Benchmark Tests
Google has finally released its long-overdue frontier model, Gemini 4 Argon. Despite being delayed and overshadowed by other models in the LLM (Large Language Model) race, Argon is still considered a competitive entrant, according to independent evaluations.
The model's performance on various benchmarks is mixed, with some areas where it leads and others where it trails behind its rivals. In knowledge work, Argon scored 68.9% on Vals Index, while GPT-6 Astra scored 63.1%. However, in automation, Argon trailed behind Claude Opus 5.5, which scored 42.5% on AutomationBenchScore.
One of the notable areas where Argon excels is in cybersecurity, scoring 68.0% on CWE-bench v1. This suggests that Google's model may be particularly effective in detecting and preventing cyber threats.
However, some Google employees have questioned Argon's performance on practical coding tasks, according to a report from Bloomberg. The company has disputed this characterization, claiming that Argon leads the frontier on a majority of benchmarks.