Astra Reaches Critical Cybersecurity Threshold as OpenAI Cools Hysteria Over AI Capabilities
OpenAI's chief scientist Jakub Pachocki recently warned against hysteria over the firm's Astra model, which has become the first in the AI race to reach the 'Critical' cybersecurity designation in its Preparedness Framework.
The Critical designation measures how much human prompting and intervention an AI model needs to find and build viable zero-day exploits across many hardened systems, or to run a full novel attack.
Astra has been touted as capable of completing all steps for full novel attacks or zero-day exploits with minimal human intervention when it launches. This is according to OpenAI's 'Path to Astra' post on September 1, which cited Astra's 100% score on ExploitBench, a test of how models can exploit known bugs.
Pachocki urged calm in the face of growing commentary that OpenAI may no longer have a hedge around the capabilities of its biggest models. He also emphasized the lab's due diligence on using chain-of-thought monitoring, which involves reading the step-by-step reasoning a model exposes as it works.