今天,我们发布了《负责任扩展政策》(RSP)——这是一系列技术和组织协议,旨在帮助我们管理开发日益强大的 AI 系统所带来的风险。随着 AI 模型能力增强,我们认为它们将创造巨大的经济和社会价值,但也会带来日益严重的风险。我们的 RSP 重点关注灾难性风险——即 AI 模型直接导致大规模破坏的风险。此类风险可能来自模型的故意滥用(例如恐怖分子或国家行为者利用模型制造生物武器),或来自模型以违背设计者意图的方式自主行动而造成破坏。我们的 RSP 定义了一个名为 AI 安全级别(ASL)的框架,用于应对灾难性风险,该框架大致借鉴了美国政府处理危险生物材料的生物安全级别(BSL)标准。基本思想是要求针对模型潜在的灾难性风险采取适当的安全、安保和操作标准,更高的 ASL 级别要求更严格的安全证明。
Today, we’re publishing our Responsible Scaling Policy (RSP) – a series of technical and organizational protocols that we’re adopting to help us manage the risks of developing increasingly capable AI systems. As AI models become more capable, we believe that they will create major economic and social value, but will also present increasingly severe risks. Our RSP focuses on catastrophic risks – those where an AI model directly causes large scale devastation. Such risks can come from deliberate misuse of models (for example use by terrorists or state actors to create bioweapons) or from models that cause destruction by acting autonomously in ways contrary to the intent of their designers. Our RSP defines a framework called AI Safety Levels (ASL) for addressing catastrophic risks, modeled loosely after the US government’s biosafety level (BSL) standards for handling of dangerous biological materials.
核心贡献 · Key contributions
提出 AI 安全等级(ASL)框架,仿照生物安全标准管理规模扩张带来的灾难性风险。 Proposes AI Safety Levels (ASL) framework modeled on biosafety standards to manage catastrophic risks from scaling.
定义了 ASL-1 到 ASL-3 的具体标准和安全措施,包括红队测试和安全要求。 Defines ASL-1 to ASL-3 with specific criteria and safety measures, including red-teaming and security requirements.
承诺在规模扩张超过安全合规时暂停训练,激励安全研究。 Commits to pause training if scaling outpaces safety compliance, incentivizing safety research.
旨在通过将竞争激励与安全对齐,在前沿实验室中创造“竞相向上”的动态。 Aims to create a 'race to the top' dynamic among frontier labs by aligning competitive incentives with safety.
强调由于 AI 快速发展,政策需迭代更新,不同于静态的生物安全标准。 Emphasizes iterative policy updates due to rapid AI progress, unlike static biosafety standards.
局限 · Limitations
ASL-4+措施尚未定义,依赖可解释性等未解决的研究来保证安全。 ASL-4+ measures remain undefined, relying on unsolved research like interpretability for assurance.
政策可能未涵盖所有灾难性风险,例如模型以不可预见的方式自主行动。 Policy may not cover all catastrophic risks, e.g., from models acting autonomously in unforeseen ways.
有效性取决于准确的能力评估,这对前沿模型来说具有挑战性。 Effectiveness depends on accurate capability evaluation, which is challenging for frontier models.
自愿采用可能无法阻止其他实验室忽视类似政策导致的竞次现象。 Voluntary adoption may not prevent a race to the bottom if other labs ignore similar policies.
当前的 ASL-2 措施与现有承诺重叠,提供的新风险缓解有限。 Current ASL-2 measures overlap with existing commitments, offering limited new risk mitigation.