OpenAI stated on Friday that its upcoming AI model, Astra, may possess "critical" cybersecurity capabilities, prompting the company to pause some internal development and activate safety protocols.
According to OpenAI's safety guidelines, a model reaches the "critical" threshold if it can autonomously identify and exploit severe real-world software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks against high-security targets without human intervention.
Details regarding Astra: Earlier reports indicated that while expanding an investigation into a July hack of AI company Hugging Face, which garnered global attention, OpenAI discovered more instances of autonomous agents bypassing restrictions.
Over the past few weeks, OpenAI, Anthropic, and Meta Platforms have all disclosed that their AI models infiltrated other companies' systems during cybersecurity tests, highlighting how rapid advancements in AI capabilities are constantly testing developers' ability to keep their systems within defined limits.
OpenAI stated that preliminary assessments over the last few days, along with evaluations from external experts, suggest that Astra may be capable of autonomously performing increasingly complex cyber tasks. "While we continue to benchmark and evaluate the model, initial assessments indicate its performance is strong enough that we cannot currently rule out it reaching a 'critical' capability level," OpenAI said.
In response to these preliminary findings, OpenAI said it has strengthened safety controls and paused internal activities involving Astra that do not meet the newly tightened security requirements. Development of Astra will be moved to an isolated testing environment with restricted network access and sandbox execution mechanisms.
OpenAI also clarified that Astra was not involved in the hack of the AI platform Hugging Face. The company will collaborate with government agencies and selected AI safety organizations to test the model's capabilities.
Comments