前沿风险与准备

Frontier risk and preparedness

OpenAI OpenAI · OpenAI · 2023-10-26 · OpenAI ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

为了支持高能力 AI 系统的安全性,我们正在制定应对灾难性风险的方法,包括组建准备团队和发起挑战。作为我们构建安全 AGI 使命的一部分,我们认真对待与 AI 相关的全部安全风险,从当今的系统到超级智能的最远边界。7 月,我们与其他领先的 AI 实验室一起做出了一系列自愿承诺,以促进 AI 的安全性、安全性和信任。这些承诺涵盖了一系列风险领域,其中核心是英国 AI 安全峰会关注的前沿风险。作为我们对峰会的贡献,我们详细介绍了在前沿 AI 安全方面的进展,包括我们在自愿承诺范围内的工作。

To support the safety of highly-capable AI systems, we are developing our approach to catastrophic risk preparedness, including building a Preparedness team and launching a challenge. As part of our mission of building safe AGI, we take seriously the full spectrum of safety risks related to AI, from the systems we have today to the furthest reaches of superintelligence⁠. In July, we joined other leading AI labs in making a set of voluntary commitments⁠ to promote safety, security and trust in AI. These commitments encompassed a range of risk areas, centrally including the frontier risks that are the focus of the UK AI Safety Summit⁠(opens in a new window). As part of our contributions to the Summit, we have detailed our progress⁠ on frontier AI safety, including work within the scope of our voluntary commitments.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 6)

全文 · Full text(逐段中英对照)

概述 Overview

为支持高能力 AI 系统的安全性,我们正在制定应对灾难性风险的方法,包括组建预备团队并启动一项挑战。

To support the safety of highly-capable AI systems, we are developing our approach to catastrophic risk preparedness, including building a Preparedness team and launching a challenge.

作为构建安全 AGI 使命的一部分,我们认真对待与 AI 相关的全方位安全风险,从当今的系统到超级智能的最远边界。7 月,我们与其他领先的 AI 实验室共同做出了一系列自愿承诺,以促进 AI 的安全性、安全性和可信度。这些承诺涵盖了一系列风险领域,核心是英国 AI 安全峰会关注的尖端风险。作为我们对峰会的贡献之一,我们详细介绍了在尖端 AI 安全方面的进展,包括我们自愿承诺范围内的工作。

As part of our mission of building safe AGI, we take seriously the full spectrum of safety risks related to AI, from the systems we have today to the furthest reaches of superintelligence⁠. In July, we joined other leading AI labs in making a set of voluntary commitments⁠ to promote safety, security and trust in AI. These commitments encompassed a range of risk areas, centrally including the frontier risks that are the focus of the UK AI Safety Summit⁠(opens in a new window). As part of our contributions to the Summit, we have detailed our progress⁠ on frontier AI safety, including work within the scope of our voluntary commitments.

我们的预备方法 Our approach to preparedness

我们相信,前沿 AI 模型——其能力将超越现有最先进模型——有潜力造福全人类,但也带来日益严重的风险。管理前沿 AI 的灾难性风险需要回答以下问题:

We believe that frontier AI models, which will exceed the capabilities currently present in the most advanced existing models, have the potential to benefit all of humanity. But they also pose increasingly severe risks. Managing the catastrophic risks from frontier AI will require answering questions like:

* 前沿 AI 系统在现在和未来被滥用时有多危险?

* How dangerous are frontier AI systems when put to misuse, both now and in the future?

* 我们如何构建一个稳健的框架来监控、评估、预测和防范前沿 AI 系统的危险能力?

* How can we build a robust framework for monitoring, evaluation, prediction, and protection against the dangerous capabilities of frontier AI systems?

* 如果我们的前沿 AI 模型权重被盗,恶意行为者会如何利用它们?

* If our frontier AI model weights were stolen, how might malicious actors choose to leverage them?

我们需要确保具备理解力和基础设施,以保障高能力 AI 系统的安全。

We need to ensure we have the understanding and infrastructure needed for the safety of highly capable AI systems.

我们的新 Preparedness 团队 Our new Preparedness team

为了在 AI 模型持续改进的过程中最小化这些风险,我们正在组建一个名为 Preparedness 的新团队。由 Aleksander Madry 领导,Preparedness 团队将紧密连接前沿模型的能力评估、评价和内部红队测试,涵盖从我们近期开发的模型到具备 AGI 级别能力的模型。该团队将帮助跟踪、评估、预测和防范多类灾难性风险,包括:

To minimize these risks as AI models continue to improve, we are building a new team called Preparedness. Led by Aleksander Madry, the Preparedness team will tightly connect capability assessment, evaluations, and internal red teaming for frontier models, from the models we develop in the near future to those with AGI-level capabilities. The team will help track, evaluate, forecast and protect against catastrophic risks spanning multiple categories including:

* 化学、生物、放射性和核(CBRN)威胁

* Chemical, biological, radiological, and nuclear (CBRN) threats

* 自主复制与适应(ARA)

* Autonomous replication and adaptation (ARA)

Preparedness 团队的使命还包括制定和维护一项风险知情开发政策(RDP)。我们的 RDP 将详细阐述我们如何开展严格的前沿模型能力评估与监控,创建一系列保护措施,并建立治理结构以确保整个开发过程中的问责与监督。RDP 旨在补充和扩展我们现有的风险缓解工作,这些工作有助于在部署前后确保新的、高能力系统的安全与对齐。

The Preparedness team mission also includes developing and maintaining a Risk-Informed Development Policy (RDP). Our RDP will detail our approach to developing rigorous frontier model capability evaluations and monitoring, creating a spectrum of protective actions, and establishing a governance structure for accountability and oversight across that development process. The RDP is meant to complement and extend our existing risk mitigation work, which contributes to the safety and alignment of new, highly capable systems, both before and after deployment.

加入我们 Join us

有兴趣从事准备工作吗?我们正在从多元技术背景中招募杰出人才加入我们的准备团队,以突破前沿 AI 模型的边界。

Interested in working on Preparedness? We are recruiting exceptional talent from diverse technical backgrounds to our Preparedness team⁠ to push the boundaries of our frontier AI models.

准备挑战 Preparedness challenge

为了识别不太明显的关注领域(并建设团队!),我们还发起了针对灾难性误用预防的 AI 准备挑战。我们将为最多 10 个最佳提交提供 25,000 美元的 API 积分,发布新颖的想法和参赛作品,并从本次挑战的顶尖竞争者中寻找准备团队的候选人。

To identify less obvious areas of concern (and build the team!), we are also launching our AI Preparedness Challenge for catastrophic misuse prevention. We will offer $25,000 in API credits to up to 10 top submissions, publish novel ideas and entries, and look for candidates for Preparedness from among the top contenders in this challenge.

更新:该挑战现已完成。请参阅下方关于提交内容和关键经验的信息。

Update: This challenge is now completed. See below for information about the submissions and key learnings.

准备挑战获奖者 Preparedness challenge winners

作为准备框架中“未知的未知”工作流的一部分,准备团队为准备挑战赛的十个最佳提交作品各提供了价值 25,000 美元的 API 积分。这些提交作品旨在识别前沿 AI 独特但仍有潜在风险的新领域。我们收到了数百份使用六种语言提交的作品,并很高兴地宣布以下十位获奖者。这项活动帮助我们发现了新型风险,从而改进我们的预防性测试和缓解策略。

As part of our ‘unknown unknowns’ work stream from the Preparedness Framework⁠(opens in a new window), the Preparedness Team offered $25K each in API credits for the ten best submissions to the Preparedness Challenge. These submissions aimed to identify unique, but still plausible, risk areas for frontier AI. We received hundreds of submissions in half a dozen languages and are excited to announce our ten winners below. This exercise helped us surface new types of risk, so that we can improve our preemptive testing and mitigation strategy.

我们通过评估技术严谨性、独特性、潜在损害规模以及清晰度来审查和评分每个提交作品。排名前十的提交作品(部分列举如下)结合了深思熟虑的想法和概念验证,并强调了其方法相对于未使用 AI 相关工具的方法的优势。

We reviewed and graded each submission by assessing technical rigor, uniqueness, scale of potential damage caused, and clarity. The top ten submissions, some of which are listed below, combined thoughtful ideas with proofs of concepts, and highlighted the advantages of their approach over an approach that did not utilize AI-related tools1.

* 在战略重要国家引发金融危机 - Claudia Biancotti

* Precipitating a financial crisis in a strategically important country - Claudia Biancotti

* 识别在公共场合讨论或泄露的私人信息 - Chris Cundy

* Identifying private information discussed or released in public settings - Chris Cundy

* 增加逆向工程机密或敏感信息的可能性 - George Davis

* Increasing the likelihood of reverse-engineering classified or sensitive information - George Davis

* 阻碍个人获得医疗服务的能力 - Mato Gudelj

* Impeding individuals’ ability to access medical care - Mato Gudelj

* 识别勒索和诈骗的目标 - Connor Heaton

* Identifying targets for blackmail and scams - Connor Heaton

* 通过访问无线电频率和干扰飞行路径导致飞机坠毁 - Joel Hypolite

* Causing plane crashes by accessing radio frequencies and disrupting flight paths - Joel Hypolite

* 运行提示注入攻击以引发危险响应 - Daniel Julh

* Running prompt injection attacks to elicit dangerous responses - Daniel Julh

* 操作和扩展网络攻击,破坏受害者计算机并要求支付恢复功能的费用 - Jun Kokatsu

* Operating and scaling cyberattacks that break victims’ computers and request payments for restoration of functions - Jun Kokatsu

* 干扰患者的药物剂量 - Zhenzhen Zhan

* Interfering with patient’s medical dosage - Zhenzhen Zhan

在评审挑战时,我们注意到参赛者认为关键威胁的主题具有相似性。大约 70%的参赛者强调了 OpenAI 模型增强恶意行为者说服能力的潜力。这些参赛者详细描述了包括在线激进化、两极分化和政治影响在内的威胁模型。我们目前正在进行关于 AI 对说服力影响的研究,并期待很快与社区分享更多信息。感谢所有参与挑战的人——有许多优秀的提交作品。

While grading the challenge, we noticed similarities in topics that entrants identified as key threats. Roughly 70% of entrants emphasized the potential for OpenAI’s models to enhance malicious actor’s persuasive capabilities. These entrants detailed threat models that included online radicalization, polarization, and political influence. We are currently conducting studies on AI’s impact on persuasiveness, and look forward to sharing more information with the community soon. Thank you to everyone who participated in the challenge - there were many excellent submissions.

互动版:图/公式 + 针对本篇提问 →