为了支持高能力 AI 系统的安全性,我们正在制定应对灾难性风险的方法,包括组建准备团队和发起挑战。作为我们构建安全 AGI 使命的一部分,我们认真对待与 AI 相关的全部安全风险,从当今的系统到超级智能的最远边界。7 月,我们与其他领先的 AI 实验室一起做出了一系列自愿承诺,以促进 AI 的安全性、安全性和信任。这些承诺涵盖了一系列风险领域,其中核心是英国 AI 安全峰会关注的前沿风险。作为我们对峰会的贡献,我们详细介绍了在前沿 AI 安全方面的进展,包括我们在自愿承诺范围内的工作。
To support the safety of highly-capable AI systems, we are developing our approach to catastrophic risk preparedness, including building a Preparedness team and launching a challenge. As part of our mission of building safe AGI, we take seriously the full spectrum of safety risks related to AI, from the systems we have today to the furthest reaches of superintelligence. In July, we joined other leading AI labs in making a set of voluntary commitments to promote safety, security and trust in AI. These commitments encompassed a range of risk areas, centrally including the frontier risks that are the focus of the UK AI Safety Summit(opens in a new window). As part of our contributions to the Summit, we have detailed our progress on frontier AI safety, including work within the scope of our voluntary commitments.
核心贡献 · Key contributions
成立了预备团队,评估和缓解前沿 AI 模型的灾难性风险。 Established a Preparedness team to evaluate and mitigate catastrophic risks from frontier AI models.
引入了基于风险的发展政策(RDP),用于能力评估和保护措施。 Introduced a Risk-Informed Development Policy (RDP) for capability evaluations and protective actions.
发起了 AI 预备挑战赛,识别未知风险领域,向最佳提交者奖励 2.5 万美元 API 积分。 Launched an AI Preparedness Challenge to identify unknown risk areas, awarding $25K API credits to top submissions.
将灾难性风险分类,包括化学、生物、放射性和核(CBRN)威胁以及自主复制与适应(ARA)。 Categorized catastrophic risks including CBRN threats and autonomous replication and adaptation (ARA).
发现约 70%的挑战提交强调 AI 增强的说服力是关键威胁。 Found that ~70% of challenge submissions highlighted AI-enhanced persuasion as a key threat.
局限 · Limitations
预备框架聚焦于前沿模型,未涵盖较不先进 AI 的风险。 The Preparedness Framework focuses on frontier models, not covering risks from less advanced AI.
挑战可能遗漏不易通过公开提交识别的风险。 The challenge may miss risks that are not easily identified through open submissions.
RDP 的有效性依赖于准确的能力预测,而这具有不确定性。 The RDP's effectiveness depends on accurate capability forecasting, which is uncertain.
该方法主要解决滥用风险,而非对齐或结构性风险。 The approach primarily addresses misuse risks, not alignment or structural risks.
自愿承诺缺乏执行机制,限制了其影响。 Voluntary commitments lack enforcement mechanisms, limiting their impact.