OpenAI published a set of proposals on Monday covering safety and safeguards for the development of cutting-edge artificial intelligence, with a central focus on alignment research and a computing technique known as "recursive self-improvement" (RSI). The company stated in a blog post that completing this transition safely requires alignment research to keep pace with these capabilities, ensuring that systems built by itself and others remain consistent with human values and under human control.
OpenAI urged international collaboration to establish frontier standards and suggested drawing on the work of existing AI safety institutions around the world. The ChatGPT developer said the technical standards should target frontier AI models and developers, as well as the management of benefits and risks associated with automated AI researchers, including RSI. The excitement around RSI stems from its potential to create foundation models capable of upgrading themselves without human intervention.
However, as RSI progresses, some technologists have raised concerns that foundation model makers could lose control over their underlying technology or fail to foresee unintended consequences, especially as AI systems grow increasingly complex and pervasive across the internet. OpenAI pointed out in the blog that fully autonomous RSI is not currently achievable and should not be pursued unless it can be done safely. Without appropriate caution and safeguards, RSI could lead to humans losing meaningful control over AI development and being unable to oversee research processes they no longer understand.
OpenAI's blog referenced the security breach at Hugging Face as a case in point. That incident did not involve RSI technology, but it served as a "dry run" for how such risks could become far more serious without robust safeguards and alignment. The debate over AI safety is heating up, with Anthropic advocating for a slowdown and growing calls for third-party evaluations.
Last week, rival company Anthropic laid out its own vision for safely developing frontier AI models, responding to a string of recent warnings from industry researchers about AI posing threats to humanity. Jacob Coxon, who previously worked at both Anthropic and OpenAI, announced his resignation roughly two weeks ago, stating that the companies were "gambling with our lives," which triggered a global conversation. In the wake of recent AI safety incidents and Coxon's public statements, Anthropic CEO Dario Amodei published a piece urging AI companies to slow down foundation model development, along with other recommendations.
Amodei also proposed the introduction of third-party evaluators within companies to audit and mitigate potential societal risks from their technology, such as exacerbating cybersecurity-related hacking or enabling biological weapons. OpenAI CEO Sam Altman, along with rival leaders including Tesla and SpaceX CEO Elon Musk, publicly backed Amodei's proposal. However, because the field of AI evaluation is still in its infancy, there is no unified consensus yet on the fundamental standards and principles required for allowing independent third parties to conduct deeper reviews of cutting-edge technology.
This has partly driven a group of AI evaluation experts to urge foundation model makers to consider a set of "minimum requirements" for enabling more thorough technical reviews and inspections, including obtaining deeper access levels and ensuring protection against retaliation for publishing unfavorable reports.
Comments