我们对对齐研究的方法

Our approach to alignment research

OpenAI OpenAI · OpenAI · 2022-08-24 · OpenAI Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们对对齐研究的方法 | OpenAI 我们正在提升 AI 系统从人类反馈中学习以及协助人类评估 AI 的能力。我们的目标是构建一个足够对齐的 AI 系统,能够帮助我们解决所有其他对齐问题。* 使用人类反馈训练 AI 系统

Our approach to alignment research | OpenAI We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems. * Training AI systems using human feedback

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 5)

全文 · Full text(逐段中英对照)

概述 Overview

我们的对齐研究方法 | OpenAI

Our approach to alignment research | OpenAI

我们正在提升 AI 系统从人类反馈中学习以及协助人类评估 AI 的能力。我们的目标是构建一个充分对齐的 AI 系统,能够帮助我们解决所有其他对齐问题。

We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.

* 使用人类反馈训练 AI 系统

* Training AI systems using human feedback

* 训练模型以协助人类评估

* Training models to assist human evaluation

* 训练 AI 系统进行对齐研究

* Training AI systems to do alignment research

* 使用人类反馈训练 AI 系统

* Training AI systems using human feedback

* 训练模型以协助人类评估

* Training models to assist human evaluation

* 训练 AI 系统进行对齐研究

* Training AI systems to do alignment research

我们的对齐研究旨在使通用人工智能(AGI)与人类价值观一致并遵循人类意图。我们采用迭代的、经验性的方法:通过尝试对齐高能力的 AI 系统,我们可以了解哪些方法有效、哪些无效,从而提升我们使 AI 系统更安全、更对齐的能力。通过科学实验,我们研究对齐技术的扩展方式及其失效点。

Our alignment research aims to make artificial general intelligence(AGI) aligned with human values and follow human intent. We take an iterative, empirical approach: by attempting to align highly capable AI systems, we can learn what works and what doesn’t, thus refining our ability to make AI systems safer and more aligned. Using scientific experiments, we study how alignment techniques scale and where they will break.

我们既在我们最强大的 AI 系统中解决对齐问题,也解决在通往 AGI 道路上预期会遇到的对齐问题。我们的主要目标是尽可能推动当前的对齐理念,并准确理解和记录它们成功或失败的原因。我们相信,即使没有根本性的新对齐理念,我们也能构建出足够对齐的 AI 系统,从而实质性地推进对齐研究本身。

We tackle alignment problems both in our most capable AI systems as well as alignment problems that we expect to encounter on our path to AGI. Our main goal is to push current alignment ideas as far as possible, and to understand and document precisely how they can succeed or why they will fail. We believe that even without fundamentally new alignment ideas, we can likely build sufficiently aligned AI systems to substantially advance alignment research itself.

未对齐的 AGI 可能对人类构成重大风险,而解决 AGI 对齐问题可能非常困难,需要全人类共同努力。因此,我们承诺在安全的情况下公开分享我们的对齐研究:我们希望透明地展示我们的对齐技术在实际中的效果,并希望每个 AGI 开发者都能使用世界上最好的对齐技术。

Unaligned AGI could pose substantial risks to humanity⁠(opens in a new window)and solving the AGI alignment problem could be so difficult that it will require all of humanity to work together. Therefore we are committed to openly sharing our alignment research when it’s safe to do so: We want to be transparent about how well our alignment techniques actually work in practice and we want every AGI developer to use the world’s best alignment techniques.

在高层次上,我们的对齐研究方法侧重于为非常智能的 AI 系统设计一个与人类意图一致的可扩展训练信号。它包含三个主要支柱:

At a high-level, our approach to alignment research focuses on engineering a scalable training signal for very smart AI systems that is aligned with human intent. It has three main pillars:

1. 使用人类反馈训练 AI 系统

1. Training AI systems using human feedback

2. 训练模型以协助人类评估

2. Training AI systems to assist human evaluation

3. 训练 AI 系统进行对齐研究

3. Training AI systems to do alignment research

使 AI 系统与人类价值观对齐也带来了一系列其他重大的社会技术挑战,例如决定这些系统应该与谁对齐。解决这些问题对于实现我们的使命很重要,但我们不在本文中讨论它们。

Aligning AI systems with human values also poses a range of other significant sociotechnical challenges, such as deciding to whom these systems should be aligned. Solving these problems is important to achievingour mission⁠, but we do not discuss them in this post.

使用人类反馈训练 AI 系统 Training AI systems using human feedback

基于人类反馈的强化学习是我们当前部署的语言模型进行对齐的主要技术。我们训练了一类名为 InstructGPT 的模型,这些模型源自预训练语言模型(如 GPT-3)。这些模型被训练以遵循人类意图:既包括指令给出的明确意图,也包括真实性、公平性和安全性等隐含意图。

RL from human feedback⁠is our main technique for aligning our deployed language models today. We train a class of models calledInstructGPT⁠(opens in a new window)derived from pretrained language models such as GPT‑3. These models are trained to follow human intent: both explicit intent given by an instruction as well as implicit intent such as truthfulness, fairness, and safety.

我们的结果表明,当前在对齐导向的微调方面存在大量唾手可得的成果:人类更偏好 InstructGPT,而非一个规模大 100 倍的预训练模型,而其微调成本不到 GPT-3 预训练算力的 2%,且需要约 20,000 小时的人类反馈。我们希望我们的工作能激励业界其他人士增加对大型语言模型对齐的投资,并提高用户对已部署模型安全性的期望。

Our results show that there is a lot of low-hanging fruit on alignment-focused fine-tuning right now: InstructGPT is preferred by humans over a 100x larger pretrained model, while its fine-tuning costs <2% of GPT‑3’s pretraining compute and about 20,000 hours of human feedback. We hope that our work inspires others in the industry to increase their investment in alignment of large language models and that it raises the bar on users’ expectations about the safety of deployed models.

我们的自然语言 API 为对齐研究提供了一个非常有用的环境:它为我们提供了丰富的反馈循环,让我们了解对齐技术在实际世界中的实际效果,这些效果基于客户愿意付费的多样化任务。平均而言,我们的客户已经更倾向于使用 InstructGPT 而非我们的预训练模型。

Our natural language API⁠is a very useful environment for our alignment research: It provides us with a rich feedback loop about how well our alignment techniques actually workin the real world⁠, grounded in a very diverse set of tasks that our customers are willing to pay money for. On average, our customers already prefer to use InstructGPT over our pretrained models.

然而,当前版本的 InstructGPT 距离完全对齐还很远:它们有时无法遵循简单指令,并非总是真实,不能可靠地拒绝有害任务,有时还会给出有偏见或有毒的回答。一些客户发现 InstructGPT 的回答明显不如预训练模型有创意,这是我们在公开基准上运行 InstructGPT 时未曾意识到的。我们也在努力发展对基于人类反馈的强化学习的更详细科学理解,以及如何提高人类反馈的质量。

Yet today’s versions of InstructGPT arequite far from fully aligned⁠: they sometimes fail to follow simple instructions, aren’t always truthful, don’t reliably refuse harmful tasks, and sometimes give biased or toxic responses. Some customers find InstructGPT’s responses significantly less creative than the pretrained models’, something we hadn’t realized from running InstructGPT on publicly available benchmarks. We are also working on developing a more detailed scientific understanding of RL from human feedback and how to improve the quality of human feedback.

对齐我们的 API 比对齐 AGI 容易得多,因为 API 上的大多数任务对人类来说并不难监督,而且我们部署的语言模型并不比人类更聪明。我们不期望基于人类反馈的强化学习足以对齐 AGI,但它是我们最感兴趣的可扩展对齐提案的核心构建块,因此完善这种方法论很有价值。

Aligning our API is much easier than aligning AGI since most tasks on our API aren’t very hard for humans to supervise and our deployed language models aren’t smarter than humans. We don’t expect RL from human feedback to be sufficient to align AGI, but it is a core building block for the scalable alignment proposals that we’re most excited about, and so it’s valuable to perfect this methodology.

训练模型以辅助人类评估 Training models to assist human evaluation

基于人类反馈的强化学习有一个根本局限:它假设人类能够准确评估我们 AI 系统正在执行的任务。如今人类在这方面相当擅长,但随着模型能力越来越强,它们将能够完成那些对人类来说更难评估的任务(例如,在大型代码库或科学论文中找出所有缺陷)。我们的模型可能会学会对人类评估者说他们想听的话,而不是告诉他们真相。为了扩展对齐,我们希望使用诸如递归奖励建模(RRM)、辩论和迭代放大等技术。

RL from human feedback has a fundamental limitation: it assumes that humans can accurately evaluate the tasks our AI systems are doing. Today humans are pretty good at this, but as models become more capable, they will be able to do tasks that are much harder for humans to evaluate (e.g., finding all the flaws in a large codebase or a scientific paper). Our models might learn to tell our human evaluators what they want to hear instead of telling them the truth. In order to scale alignment, we want to use techniques likerecursive reward modeling (RRM)⁠(opens in a new window),debate⁠, anditerated amplification⁠.

目前我们的主要方向基于 RRM:我们训练模型,使其能够在人类难以直接评估的任务上辅助人类评估我们的模型。例如:

Currently our main direction is based on RRM: we train models that can assist humans at evaluating our models on tasks that are too difficult for humans to evaluate directly. For example:

* 我们训练了一个模型来总结书籍。如果人类不熟悉书籍内容,评估书籍摘要需要很长时间,但我们的模型可以通过撰写章节摘要来辅助人类评估。

* We trained a model tosummarize books⁠. Evaluating book summaries takes a long time for humans if they are unfamiliar with the book, but our model can assist human evaluation by writing chapter summaries.

* 我们训练了一个模型,通过浏览网页并提供引用和链接来辅助人类评估事实准确性。在简单问题上,该模型的输出已经比人类撰写的回答更受青睐。

* We trained a model toassist humans at evaluating the factual accuracy⁠by browsing the web and providing quotes and links. On simple questions, this model’s outputs are already preferred to responses written by humans.

* 我们训练了一个模型对其自身输出撰写批评性评论:在基于查询的摘要任务中,借助批评性评论,人类在模型输出中发现的缺陷平均增加了 50%。即使我们要求人类撰写看似合理但不正确的摘要,这一结果仍然成立。

* We trained a model towrite critical comments on its own outputs⁠: On a query-based summarization task, assistance with critical comments increases the flaws humans find in model outputs by 50% on average. This holds even if we ask humans to write plausible looking but incorrect summaries.

* 我们正在创建一组编码任务,这些任务被选为对于未受辅助的人类来说极难可靠评估。我们希望很快发布这个数据集。

* We are creating a set of coding tasks selected to be very difficult to evaluate reliably for unassisted humans. We hope to release this data set soon.

我们的对齐技术需要即使在我们 AI 系统提出非常有创意的解决方案(如 AlphaGo 的第 37 手)时也能发挥作用,因此我们特别感兴趣的是训练模型来辅助人类区分正确与误导性或欺骗性的解决方案。我们相信,要尽可能多地了解如何使 AI 辅助评估在实践中发挥作用,最好的方法就是构建 AI 助手。

Our alignment techniques need to work even if our AI systems are proposing very creative solutions (likeAlphaGo’s move 37⁠(opens in a new window)), thus we are especially interested in training models to assist humans to distinguish correct from misleading or deceptive solutions. We believe the best way to learn as much as possible about how to make AI-assisted evaluation work in practice is to build AI assistants.

训练 AI 系统进行对齐研究 Training AI systems to do alignment research

目前尚无已知的无限可扩展的对齐问题解决方案。随着 AI 的持续进步,我们预计会遇到一系列当前系统中尚未观察到的新对齐问题。其中一些问题我们目前已有预见,而另一些则将是全新的。

There is currently no known indefinitely scalable solution to the alignment problem. As AI progress continues, we expect to encounter a number of new alignment problems that we don’t observe yet in current systems. Some of these problems we anticipate now and some of them will be entirely new.

我们认为找到无限可扩展的解决方案可能非常困难。相反,我们采取更务实的方法:构建并对齐一个能够比人类更快、更好地进行对齐研究的系统。

We believe that finding an indefinitely scalable solution is likely very difficult. Instead, we aim for a more pragmatic approach: building and aligning a system that can make faster and better alignment research progress than humans can.

随着我们在这方面取得进展,我们的 AI 系统可以接管越来越多的对齐工作,并最终构思、实现、研究和开发出比我们现有技术更好的对齐技术。它们将与人类合作,确保它们自己的后继者与人类更加对齐。

As we make progress on this, our AI systems can take over more and more of our alignment work and ultimately conceive, implement, study, and develop better alignment techniques than we have now. They will work together with humans to ensure that their own successors are more aligned with humans.

我们相信,评估对齐研究比产生对齐研究要容易得多,尤其是在有评估辅助的情况下。因此,人类研究人员将把越来越多的精力集中在审查 AI 系统完成的对齐研究上,而不是自己进行这些研究。我们的目标是训练模型使其如此对齐,以至于我们可以将几乎所有的对齐研究所需的认知劳动都外包出去。

We believe that evaluating alignment research is substantially easier than producing it, especially when provided with evaluation assistance. Therefore human researchers will focus more and more of their effort on reviewing alignment research done by AI systems instead of generating this research by themselves. Our goal is to train models to be so aligned that we can off-load almost all of the cognitive labor required for alignment research.

重要的是,我们只需要在相关领域具有人类水平能力的“较窄”AI 系统,就能在对齐研究上做得和人类一样好。我们预计这些 AI 系统比通用系统或比人类聪明得多的系统更容易对齐。

Importantly, we only need “narrower” AI systems that have human-level capabilities in the relevant domains to do as well as humans on alignment research. We expect these AI systems are easier to align than general-purpose systems or systems much smarter than humans.

语言模型特别适合自动化对齐研究,因为它们通过阅读互联网“预装”了大量关于人类价值观的知识和信息。开箱即用时,它们不是独立的智能体,因此不会在世界上追求自己的目标。进行对齐研究时,它们不需要无限制地访问互联网。然而,许多对齐研究任务可以表述为自然语言或编码任务。

Language models are particularly well-suited for automating alignment research because they come “preloaded” with a lot of knowledge and information about human values from reading the internet. Out of the box, they aren’t independent agents and thus don’t pursue their own goals in the world. To do alignment research they don’t need unrestricted access to the internet. Yet a lot of alignment research tasks can be phrased as natural language or coding tasks.

未来版本的 WebGPT、InstructGPT 和 Codex 可以作为对齐研究助手的基础,但它们目前的能力还不够。虽然我们不知道我们的模型何时才能足够有能力为对齐研究做出有意义的贡献,但我们认为提前开始很重要。一旦我们训练出一个可能有用的模型,我们计划让它可供外部对齐研究社区使用。

Future versions ofWebGPT⁠,InstructGPT⁠, andCodex⁠can provide a foundation as alignment research assistants, but they aren’t sufficiently capable yet. While we don’t know when our models will be capable enough to meaningfully contribute to alignment research, we think it’s important to get started ahead of time. Once we train a model that could be useful, we plan to make it accessible to the external alignment research community.

局限性 Limitations

我们对这种对齐 AGI 的方法感到非常兴奋,但我们预计随着对 AI 技术发展的了解加深,它需要被调整和改进。我们的方法也存在一些重要的局限性:

We’re very excited about this approach towards aligning AGI, but we expect that it needs to be adapted and improved as we learn more about how AI technology develops. Our approach also has a number of important limitations:

* 这里提出的路径低估了鲁棒性和可解释性研究的重要性,OpenAI 目前在这两个领域的投入不足。如果这符合你的背景,请申请我们的研究科学家职位!

* The path laid out here underemphasizes the importance of robustness and interpretability research, two areas OpenAI is currently underinvested in. If this fits your profile, please apply for our research scientist positions!

* 使用 AI 辅助进行评估有可能放大甚至加剧 AI 助手中存在的细微不一致、偏见或漏洞。

* Using AI assistance for evaluation has the potential to scale up or amplify even subtle inconsistencies, biases, or vulnerabilities present in the AI assistant.

* 对齐 AGI 可能涉及与对齐当今 AI 系统截然不同的问题。我们预计这一转变会是某种程度上的连续过程,但如果存在重大的不连续性或范式转变,那么从对齐 InstructGPT 等模型中学到的大部分经验可能不会直接有用。

* Aligning AGI likely involves solving very different problems than aligning today’s AI systems. We expect the transition to be somewhat continuous, but if there are major discontinuities or paradigm shifts, then most lessons learned from aligning models like InstructGPT might not be directly useful.

* 对齐问题最困难的部分可能不在于为我们的 AI 系统设计可扩展且对齐的训练信号。即使如此,这样的训练信号也是必要的。

* The hardest parts of the alignment problem might not be related to engineering a scalable and aligned training signal for our AI systems. Even if this is true, such a training signal will be necessary.

* 对齐能够有意义地加速对齐研究的模型,可能并不比对齐 AGI 本质上更容易。换句话说,如果未能正确对齐,那些能够帮助对齐研究的最弱模型可能已经过于危险。如果这是真的,我们将无法从自己的系统中获得太多帮助来解决对齐问题。

* It might not be fundamentally easier to align models that can meaningfully accelerate alignment research than it is to align AGI. In other words, the least capable models that can help with alignment research might already be too dangerous if not properly aligned. If this is true, we won’t get much help from our own systems for solving alignment problems.

我们正在为这一研究方向招聘更多优秀人才!如果你对此感兴趣,我们正在招聘研究工程师⁠(在新窗口中打开)和研究科学家⁠(在新窗口中打开)。

_We’re looking to hire more talented people for this line of research! If this interests you, we’re hiringResearch Engineers_⁠(opens in a new window)andResearch Scientists⁠(opens in a new window).

互动版:图/公式 + 针对本篇提问 →