Claim of 98.6% Score for OpenAI’s GPT-6 Astra Exposed as Baseless
A claim has been circulating on social media that OpenAI's GPT-6 Astra scored an astonishing 98.6% on the ARC-AGI-3 benchmark, which would be a significant leap in AI capability. However, upon closer inspection, this claim doesn't hold up to scrutiny.
The ARC-AGI-3 benchmark, launched on March 25, 2026, tests whether AI models can navigate interactive environments without instructions or predefined objectives. Initially, scores came in below 1%, indicating just how challenging the benchmark is.
Claude Opus 5 from Anthropic currently leads the leaderboard with a score of around 30.2%. OpenAI's GPT-5.6 Sol, its most recent model with verified results, scored 7.78% under official testing conditions and 38.3% when using custom settings.
No confirmed scores exist for Astra on ARC-AGI-3, and as of early September 2026, OpenAI hasn't announced an official release date or branding for GPT-6. The company did preview Astra in August 2026, where it demonstrated solving around 10 open math problems that had stumped mathematicians.