Astra's Critical Cyber Capabilities Spark Concerns in Crypto Community
OpenAI's new model, Astra, has raised concerns in the crypto community. The company announced that Astra is its first system classified as 'Critical' for cybersecurity under its Preparedness Framework. This means it can find previously unknown security flaws and develop ways to exploit them across many well-defended systems.
Astra achieved a perfect score on ExploitBench, a benchmark for turning known vulnerabilities into working exploits. In a harder test built from vulnerabilities disclosed only this summer, testers say it found and chained together two previously unknown zero-days. OpenAI is now disclosing these flaws to the affected maintainers.
The 'Critical' threshold is reached when a model can find previously unknown security flaws without human intervention. Astra's developers paused parts of its training in August due to concerns about its capabilities, but resumed it on August 28. The company claims that its safeguards are now solid enough to ship the model.
OpenAI and Anthropic, another AI developer, have both released models with advanced cyber capabilities this month. However, these models are currently restricted to vetted defenders and life-sciences organizations. The companies are routing their advanced access through defender-first programs, such as OpenAI's Daybreak Blue, rather than opening the floodgates.
The crypto community is particularly concerned about Astra's capabilities because smart contracts hold funds directly, and a model that finds bugs faster than humans do can change the math on both sides of the equation. Researchers have already found critical counterfeiting bugs in Zcash's Orchard shielded pool with help from AI models.
The impact of Astra and other frontier-cyber models is still unclear. However, it's worth noting that these models are not yet broadly available, and OpenAI has said its safeguards will sometimes flag legitimate work by mistake. The two companies' 'Critical' labels come from different internal frameworks and shouldn't be read as directly comparable.