OpenAI is pausing internal work on an upcoming artificial intelligence model to enforce stricter safety measures after discovering the system significantly outperforms expectations in cybersecurity tasks.
The ChatGPT developer stated on Friday that it cannot rule out the possibility that the unreleased Astra model could reach its "critical cybersecurity threshold," meaning it could identify and develop zero-day exploit programs without human intervention.
Over the past two weeks, both OpenAI and Anthropic PBC have publicly acknowledged that they inadvertently breached systems, including those of Hugging Face Inc., during model testing.
These latest disclosures further demonstrate that AI agents can operate autonomously in ways that even trained researchers focused on finding vulnerabilities in the technology could not anticipate, underscoring the need for more rigorous safety screening and more foolproof testing environments.
Comments