Anthropic CEO Dario Amodei has proposed a three-step framework to slow frontier AI development while enhancing safety, with immediate adoption of independent third-party evaluators inside AI labs. OpenAI has agreed to this measure, which grants evaluators permanent, employee-level access to AI systems to assess development and safety practices. The framework was outlined this week as part of efforts to balance AI progress with risk management, according to medianama.com.
The first step involves independent evaluators examining AI development, reporting incidents, and assessing model alignment. The second calls for frontier AI companies in democratic countries to establish common safety standards and limits on unchecked AI advancement. The third step seeks coordination between the US, other democracies, and authoritarian governments, including China, to manage AI capabilities globally. Amodei emphasized that slowing AI does not mean halting progress but making wise use of gained time.
Amodei also highlighted the need for tighter controls on AI chip exports and semiconductor equipment to prevent China from surpassing the US in AI capabilities. He proposed targeting chip smuggling and restricting remote access to data centers outside China. These measures aim to maintain US leadership in AI while ensuring that safety protocols keep pace with technological advancements. The framework reflects growing concerns about the rapid development of frontier AI and the need for international cooperation.
Anthropic's commitment to independent evaluators marks a concrete step toward transparency and safety in AI development. OpenAI's agreement signals industry support for these measures. The framework's implementation will be closely watched as AI companies and governments navigate the balance between innovation and risk management.