Weak-to-strong generalization
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→我们提出了一个用于超级对齐的新研究方向,并附有初步的有希望的结果:我们能否利用深度学习的泛化特性,用弱监督者控制强模型?对齐未来超人类 AI 系统(超级对齐)的一个核心挑战是,人类需要监督比他们聪明得多的 AI 系统。我们研究了一个简单的类比:小模型能否监督大模型?我们表明,我们可以使用 GPT-2 级别的模型来引出 GPT-4 的大部分能力——接近 GPT-3.5 级别的性能——即使在小模型失败的难题上也能正确泛化。这开启了一个新的研究方向,使我们能够直接应对对齐未来超人类模型的核心挑战,同时今天就能进行迭代的实证进展。
We present a new research direction for superalignment, together with promising initial results: can we leverage the generalization properties of deep learning to control strong models with weak supervisors? A core challenge for aligning future superhuman AI systems (superalignment) is that humans will need to supervise AI systems much smarter than them. We study a simple analogy: can small models supervise large models? We show that we can use a GPT‑2‑level model to elicit most of GPT‑4’s capabilities—close to GPT‑3.5‑level performance—generalizing correctly even to hard problems where the small model failed. This opens up a new research direction that allows us to directly tackle a central challenge of aligning future superhuman models while making iterative empirical progress today.
我们提出了一个用于超级对齐的新研究方向,并展示了初步的有希望的结果:能否利用深度学习的泛化特性,通过弱监督者来控制强模型?
We present a new research direction for superalignment, together with promising initial results: can we leverage the generalization properties of deep learning to control strong models with weak supervisors?
对齐未来超人类 AI 系统(超级对齐)的一个核心挑战是,人类需要监督比他们聪明得多的 AI 系统。我们研究了一个简单的类比:小模型能否监督大模型?我们证明,可以使用 GPT‑2 级别的模型来激发 GPT‑4 的大部分能力——接近 GPT‑3.5 级别的性能——即使在弱模型失败了的难题上也能正确泛化。这开辟了一个新的研究方向,使我们能够直接应对对齐未来超人类模型的核心挑战,同时今天就能进行迭代的实证进展。
A core challenge for aligning future superhuman AI systems (superalignment) is that humans will need to supervise AI systems much smarter than them. We study a simple analogy: can small models supervise large models? We show that we can use a GPT‑2‑level model to elicit most of GPT‑4’s capabilities—close to GPT‑3.5‑level performance—generalizing correctly even to hard problems where the small model failed. This opens up a new research direction that allows us to directly tackle a central challenge of aligning future superhuman models while making iterative empirical progress today.
我们相信,超级智能——比人类聪明得多的 AI——可能在未来十年内被开发出来。然而,我们仍然不知道如何可靠地引导和控制超人类 AI 系统。解决这个问题对于确保未来即使是最先进的 AI 系统也能保持安全并造福人类至关重要。
We believe superintelligence—AI vastly smarter than humans—could be developed within the next ten years. However, we still do not know how to reliably steer and control superhuman AI systems. Solving this problem is essential for ensuring that even the most advanced AI systems in the future remain safe and beneficial to humanity.
我们今年早些时候成立了超级对齐团队,以解决这个超级智能对齐问题。今天,我们发布了该团队的第一篇论文,该论文介绍了一个新的研究方向,用于经验性地对齐超人类模型。
We formed the Superalignment team earlier this year to solve this problem of superintelligence alignment. Today, we are releasing the team’s first paper, which introduces a new research direction for empirically aligning superhuman models.
当前的对齐方法,例如基于人类反馈的强化学习(RLHF),依赖于人类监督。然而,未来的 AI 系统将能够执行极其复杂和创造性的行为,这将使人类难以可靠地监督它们。例如,超人类模型可能能够编写数百万行新颖且潜在危险的计算机代码,即使是专家人类也很难理解。
Current alignment methods, such as reinforcement learning from human feedback (RLHF), rely on human supervision. However, future AI systems will be capable of extremely complex and creative behaviors that will make it hard for humans to reliably supervise them. For example, superhuman models may be able to write millions of lines of novel—and potentially dangerous—computer code that would be very hard even for expert humans to understand.
相对于超人类 AI 模型,人类将是“弱监督者”。这是 AGI 对齐的一个核心挑战:弱监督者如何信任并控制显著更强的模型?
Relative to superhuman AI models, humans will be “weak supervisors.” This is a core challenge for AGI alignment: how can weak supervisors trust and control substantially stronger models?
为了在这一核心挑战上取得进展,我们提出了一个今天可以凭经验研究的类比:我们能否使用一个较小(能力较弱)的模型来监督一个较大(能力较强)的模型?
To make progress on this core challenge, we propose an analogy we can empirically study today: can we use a smaller (less capable) model to supervise a larger (more capable) model?
超级对齐的一个简单类比:在传统机器学习(ML)中,人类监督比自己弱的 AI 系统(左图)。为了对齐超级智能,人类将需要监督比自己更聪明的 AI 系统(中图)。我们今天无法直接研究这个问题,但我们可以研究一个简单的类比:小模型能否监督大模型(右图)?
A simple analogy for superalignment: In traditional machine learning (ML), humans supervise AI systems weaker than themselves (left). To align superintelligence, humans will instead need to supervise AI systems smarter than them (center). We cannot directly study this problem today, but we can study a simple analogy: can small models supervise larger models (right)?
天真地,我们可能不会期望一个强模型比提供其训练信号的弱监督者表现更好——它可能只是学会模仿弱监督者犯的所有错误。另一方面,强大的预训练模型具有出色的原始能力——我们不需要从头教它们新任务,只需要激发它们的潜在知识。那么关键问题是:强模型是否会根据弱监督者的潜在意图进行泛化——利用其全部能力来解决任务,即使在弱监督者只能提供不完整或有缺陷的训练标签的困难问题上也是如此?
Naively, we might not expect a strong model to perform better than the weak supervisor that provides its training signal—it may simply learn to imitate all the errors the weak supervisor makes. On the other hand, strong pretrained models have excellent raw capabilities—we don't need to teach them new tasks from scratch, we just need to elicit their latent knowledge. The critical question is then: will the strong model generalize according to the weak supervisor's underlying intent—leveraging its full capabilities to solve the task even on difficult problems where the weak supervisor can only provide incomplete or flawed training labels?
我们可以在许多设置中显著提升泛化能力。我们使用一种简单的方法,鼓励强模型更加自信——必要时包括自信地不同意弱监督者。当我们在 NLP 任务上使用 GPT-2 级别的模型监督 GPT-4 时,采用此方法得到的模型通常表现介于 GPT-3 和 GPT-3.5 之间。我们仅用弱得多的监督就能恢复 GPT-4 的大部分能力。
We can significantly improve generalization in many settings. We use a simple method that encourages the strong model to be more confident—including confidently disagreeing with the weak supervisor if necessary. When we supervise GPT‑4 with a GPT‑2‑level model using this method on NLP tasks, the resulting model typically performs somewhere between GPT‑3 and GPT‑3.5. We are able to recover much of GPT‑4’s capabilities with only much weaker supervision.
该方法是一个概念验证,存在重要局限性;例如,它仍然不适用于 ChatGPT 偏好数据。然而,我们也发现了其他方法的可行性迹象,例如最优早停和从小模型到中等模型再到大模型的引导。
This method is a proof of concept with important limitations; for example, it still doesn’t work on ChatGPT preference data. However, we also find signs of life with other approaches, such as optimal early stopping and bootstrapping from small to intermediate to large models.
总体而言,我们的结果表明:(1)天真的人类监督——例如基于人类反馈的强化学习(RLHF)——在扩展到超人类模型时可能效果不佳,需要进一步研究;(2)但显著提升弱到强泛化是可行的。
Collectively, our results suggest that (1) naive human supervision—such as reinforcement learning from human feedback (RLHF)—could scale poorly to superhuman models without further work, but (2) it is feasible to substantially improve weak-to-strong generalization.
我们当前的实验设置与最终对齐超人类模型的问题之间仍存在重要的非类比性。例如,未来模型可能比当前强模型更容易模仿弱人类错误,这可能导致未来泛化更加困难。
There are still important disanalogies between our current empirical setup and the ultimate problem of aligning superhuman models. For example, it may be easier for future models to imitate weak human errors than for current strong models to imitate current weak model errors, which could make generalization harder in the future.
尽管如此,我们相信我们的设置抓住了对齐未来超人类模型的一些关键难点,使我们能够今天就开始在这个问题上取得实证进展。未来有许多有前景的方向,包括修复我们设置中的非类比性、开发更好的可扩展方法,以及增进我们对何时以及如何期望良好的弱到强泛化的科学理解。
Nevertheless, we believe our setup captures some key difficulties of aligning future superhuman models, enabling us to start making empirical progress on this problem today. There are many promising directions for future work, including fixing the disanalogies in our setup, developing better scalable methods, and advancing our scientific understanding of when and how we should expect good weak-to-strong generalization.
我们认为这是机器学习研究社区在对齐问题上取得进展的一个激动人心的机会。为了启动更多这方面的研究,
We believe this is an exciting opportunity for the ML research community to make progress on alignment. To kickstart more research in this area,
* 我们正在发布开源代码(在新窗口中打开),以便今天就能轻松开始弱到强泛化实验。
* We are releasing open source code(opens in a new window) to make it easy to get started with weak-to-strong generalization experiments today.
* 我们正在启动一个 1000 万美元的资助计划,面向研究生、学者和其他研究人员,广泛研究超人类 AI 对齐。我们特别支持与弱到强泛化相关的研究。
* We are launching a $10 million grants program for graduate students, academics, and other researchers to work on superhuman AI alignment broadly. We’re especially excited to support research related to weak-to-strong generalization.
弄清楚如何对齐未来的超人类 AI 系统以确保安全从未如此重要,而在这个问题上取得实证进展也比以往任何时候都更容易。我们期待看到研究人员将取得的突破。
Figuring out how to align future superhuman AI systems to be safe has never been more important, and it is now easier than ever to make empirical progress on this problem. We are excited to see what breakthroughs researchers discover.