Anthropic 负责任扩展政策发布

Announcing Anthropic's Responsible Scaling Policy

Anthropic Anthropic · Anthropic · 2023-09-19 · Anthropic ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

今天,我们发布了《负责任扩展政策》(RSP)——这是一系列技术和组织协议,旨在帮助我们管理开发日益强大的 AI 系统所带来的风险。随着 AI 模型能力增强,我们认为它们将创造巨大的经济和社会价值,但也会带来日益严重的风险。我们的 RSP 重点关注灾难性风险——即 AI 模型直接导致大规模破坏的风险。此类风险可能来自模型的故意滥用(例如恐怖分子或国家行为者利用模型制造生物武器),或来自模型以违背设计者意图的方式自主行动而造成破坏。我们的 RSP 定义了一个名为 AI 安全级别(ASL)的框架,用于应对灾难性风险,该框架大致借鉴了美国政府处理危险生物材料的生物安全级别(BSL)标准。基本思想是要求针对模型潜在的灾难性风险采取适当的安全、安保和操作标准,更高的 ASL 级别要求更严格的安全证明。

Today, we’re publishing our Responsible Scaling Policy (RSP) – a series of technical and organizational protocols that we’re adopting to help us manage the risks of developing increasingly capable AI systems. As AI models become more capable, we believe that they will create major economic and social value, but will also present increasingly severe risks. Our RSP focuses on catastrophic risks – those where an AI model directly causes large scale devastation. Such risks can come from deliberate misuse of models (for example use by terrorists or state actors to create bioweapons) or from models that cause destruction by acting autonomously in ways contrary to the intent of their designers. Our RSP defines a framework called AI Safety Levels (ASL) for addressing catastrophic risks, modeled loosely after the US government’s biosafety level (BSL) standards for handling of dangerous biological materials.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

全文 · Full text(逐段中英对照)

介绍 Anthropic 的负责任扩展政策 Introducing Anthropic's Responsible Scaling Policy

今天,我们发布《负责任扩展政策》(RSP)——一系列我们正在采用的技术和组织协议,以帮助管理开发日益强大 AI 系统的风险。

Today, we’re publishing our Responsible Scaling Policy (RSP) – a series of technical and organizational protocols that we’re adopting to help us manage the risks of developing increasingly capable AI systems.

随着 AI 模型变得更加强大,我们相信它们将创造巨大的经济和社会价值,但也将带来日益严重的风险。我们的 RSP 聚焦于灾难性风险——即 AI 模型直接导致大规模破坏的风险。此类风险可能来自模型的故意滥用(例如恐怖分子或国家行为者利用模型制造生物武器),或来自模型以违背设计者意图的方式自主行动而造成破坏。

As AI models become more capable, we believe that they will create major economic and social value, but will also present increasingly severe risks. Our RSP focuses on catastrophic risks – those where an AI model directly causes large scale devastation. Such risks can come from deliberate misuse of models (for example use by terrorists or state actors to create bioweapons) or from models that cause destruction by acting autonomously in ways contrary to the intent of their designers.

我们的 RSP 定义了一个名为 AI 安全等级(ASL)的框架,用于应对灾难性风险,该框架大致借鉴了美国政府处理危险生物材料的生物安全等级(BSL)标准。基本思想是要求与模型潜在灾难性风险相适应的安全、安保和操作标准,更高的 ASL 等级要求更严格的安全证明。

Our RSP defines a framework called AI Safety Levels (ASL) for addressing catastrophic risks, modeled loosely after the US government’s biosafety level (BSL) standards for handling of dangerous biological materials. The basic idea is to require safety, security, and operational standards appropriate to a model’s potential for catastrophic risk, with higher ASL levels requiring increasingly strict demonstrations of safety.

ASL 系统的简要总结如下:

A very abbreviated summary of the ASL system is as follows:

* ASL-1 指不构成重大灾难性风险的系统,例如 2018 年的 LLM 或仅下棋的 AI 系统。

* ASL-1 refers to systems which pose no meaningful catastrophic risk, for example a 2018 LLM or an AI system that only plays chess.

* ASL-2 指显示出危险能力早期迹象的系统——例如能够提供如何制造生物武器的指令——但由于可靠性不足或提供的信息并非搜索引擎无法提供,这些信息尚不实用。当前的 LLM,包括 Claude,似乎属于 ASL-2。

* ASL-2 refers to systems that show early signs of dangerous capabilities – for example ability to give instructions on how to build bioweapons – but where the information is not yet useful due to insufficient reliability or not providing information that e.g. a search engine couldn’t. Current LLMs, including Claude, appear to be ASL-2.

* ASL-3 指相比非 AI 基线(如搜索引擎或教科书)显著增加灾难性滥用风险的系统,或显示出低级自主能力的系统。

* ASL-3 refers to systems that substantially increase the risk of catastrophic misuse compared to non-AI baselines (e.g. search engines or textbooks) OR that show low-level autonomous capabilities.

* ASL-4 及更高等级(ASL-5+)尚未定义,因为距离当前系统太远,但可能涉及灾难性滥用潜力和自主性的质的升级。

* ASL-4 and higher (ASL-5+) is not yet defined as it is too far from present systems, but will likely involve qualitative escalations in catastrophic misuse potential and autonomy.

每个 ASL 等级的定义、标准和安全措施在主文档中有详细描述,但总体而言,ASL-2 措施代表我们当前的安全和安保标准,与我们近期的白宫承诺高度重叠。ASL-3 措施包括更严格的标准,需要大量的研究和工程努力才能及时遵守,例如异常强大的安全要求,以及承诺如果 ASL-3 模型在世界级红队对抗测试中显示出任何有意义的灾难性滥用风险,则不部署该模型(这与仅承诺进行红队测试形成对比)。我们的 ASL-4 措施尚未制定(我们的承诺是在达到 ASL-3 之前制定它们),但可能需要当前尚未解决的研究问题的方法来保证,例如使用可解释性方法从机制上证明模型不太可能参与某些灾难性行为。

The definition, criteria, and safety measures for each ASL level are described in detail in the main document, but at a high level, ASL-2 measures represent our current safety and security standards and overlap significantly with our recent White House commitments. ASL-3 measures include stricter standards that will require intense research and engineering effort to comply with in time, such as unusually strong security requirements and a commitment not to deploy ASL-3 models if they show any meaningful catastrophic misuse risk under adversarial testing by world-class red-teamers (this is in contrast to merely a commitment to perform red-teaming). Our ASL-4 measures aren’t yet written (our commitment is to write them before we reach ASL-3), but may require methods of assurance that are unsolved research problems today, such as using interpretability methods to demonstrate mechanistically that a model is unlikely to engage in certain catastrophic behaviors.

我们设计 ASL 系统旨在有效应对灾难性风险与激励有益应用和安全进展之间取得平衡。一方面,ASL 系统隐含地要求,如果我们的 AI 扩展速度超过我们遵守必要安全程序的能力,则暂时暂停更强大模型的训练。但它以一种直接激励我们解决必要安全问题以解锁进一步扩展的方式做到这一点,并允许我们使用前一个 ASL 等级中最强大的模型作为开发下一等级安全特性的工具。1 如果前沿实验室将其作为标准采用,我们希望这能创造一种“竞相向上”的动态,其中竞争性激励直接用于解决安全问题。

We have designed the ASL system to strike a balance between effectively targeting catastrophic risk and incentivising beneficial applications and safety progress. On the one hand, the ASL system implicitly requires us to temporarily pause training of more powerful models if our AI scaling outstrips our ability to comply with the necessary safety procedures. But it does so in a way that directly incentivizes us to solve the necessary safety issues as a way to unlock further scaling, and allows us to use the most powerful models from the previous ASL level as a tool for developing safety features for the next level.1 If adopted as a standard across frontier labs, we hope this might create a “race to the top” dynamic where competitive incentives are directly channeled into solving safety problems.

从商业角度来看,我们希望明确,我们的 RSP 不会改变 Claude 的当前用途或破坏我们产品的可用性。相反,它应被视为类似于汽车或航空工业中进行的上市前测试和安全特性设计,其目标是在产品投放市场之前严格证明其安全性,这最终使客户受益。

From a business perspective, we want to be clear that our RSP will not alter current uses of Claude or disrupt availability of our products. Rather, it should be seen as analogous to the pre-market testing and safety feature design conducted in the automotive or aviation industry, where the goal is to rigorously demonstrate the safety of a product before it is released onto the market, which ultimately benefits customers.

Anthropic 的 RSP 已由其董事会正式批准,变更必须经董事会与长期利益信托协商后批准。在完整文档中,我们描述了许多程序性保障措施,以确保评估过程的完整性。

Anthropic’s RSP has been formally approved by its board and changes must be approved by the board following consultations with the Long Term Benefit Trust. In the full document we describe a number of procedural safeguards to ensure the integrity of the evaluation process.

然而,我们想强调的是,这些承诺是我们当前的最佳猜测,是一个早期迭代,我们将在此基础上继续发展。AI 领域的快速发展和诸多不确定性意味着,与相对稳定的 BSL 系统不同,快速迭代和路线修正几乎肯定是必要的。

However, we want to emphasize that these commitments are our current best guess, and an early iteration that we will build on. The fast pace and many uncertainties of AI as a field imply that, unlike the relatively stable BSL system, rapid iteration and course correction will almost certainly be necessary.

完整文档可在此处阅读。我们希望它能对政策制定者、第三方非营利组织以及其他面临类似部署决策的公司提供有用的启发。

The full document can be read here. We hope that it provides useful inspiration to policymakers, third party nonprofit organizations, and other companies facing similar deployment decisions.

我们感谢 ARC Evals 在支持我们 RSP 承诺开发方面的关键见解和专业知识,特别是在自主能力评估方面。我们发现他们在 AI 风险评估方面的专业知识在我们设计评估程序时起到了关键作用。我们也认可 ARC Evals 在发起和引领其更广泛的 ARC 负责任扩展政策框架开发方面的领导作用,该框架启发了我们的方法。

_We thank ARC Evals for their key insights and expertise supporting the development of our RSP commitments, particularly regarding evaluations for autonomous capabilities. We found their expertise in AI risk assessment to be instrumental as we designed our evaluation procedures. We also recognize ARC Evals' leadership in originating and spearheading the development of their broader ARC Responsible Scaling Policy framework, which inspired our approach._

1. 总体而言,Anthropic 一直发现,与前沿 AI 模型合作是开发减轻 AI 风险新方法的关键要素。

1. As a general matter, Anthropic has consistently found that working with frontier AI models is an essential ingredient in developing new methods to mitigate the risk of AI.

互动版:图/公式 + 针对本篇提问 →