OpenAI's Astra Hits Critical Cybersecurity Threshold with Robust Safeguards
OpenAI's Astra has achieved critical cybersecurity capability under its Preparedness Framework, making it the first AI model to do so. This milestone demonstrates Astra's ability to autonomously identify and exploit zero-day vulnerabilities in hardened systems, as well as execute sophisticated cyberattack strategies without human intervention.
Astra's capabilities outstrip those of its predecessor, GPT-5.6 Sol, by achieving perfect scores on ExploitBench and uncovering two previously unknown zero-day vulnerabilities during testing. These vulnerabilities have since been disclosed to the relevant maintainers.
OpenAI delayed Astra's development and release to strengthen protections against cyber misuse and unauthorized actions. Enhanced safeguards include training the model to deny harmful requests, implementing stricter monitoring systems, and improving defenses against jailbreaks. Astra reportedly refuses 91.5% of disallowed cyber assistance requests, a significant improvement from GPT-5.6 Sol's 59% refusal rate.
Access to Astra's most advanced cybersecurity features will initially be limited. An alpha testing group will gain early access, with broader availability through OpenAI's Daybreak Blue program to support defensive applications. OpenAI has also committed to publishing detailed evaluations and safeguards in Astra's system card upon release.