我们提出了一个用于超级对齐的新研究方向,并附有初步的有希望的结果:我们能否利用深度学习的泛化特性,用弱监督者控制强模型?对齐未来超人类 AI 系统(超级对齐)的一个核心挑战是,人类需要监督比他们聪明得多的 AI 系统。我们研究了一个简单的类比:小模型能否监督大模型?我们表明,我们可以使用 GPT-2 级别的模型来引出 GPT-4 的大部分能力——接近 GPT-3.5 级别的性能——即使在小模型失败的难题上也能正确泛化。这开启了一个新的研究方向,使我们能够直接应对对齐未来超人类模型的核心挑战,同时今天就能进行迭代的实证进展。
We present a new research direction for superalignment, together with promising initial results: can we leverage the generalization properties of deep learning to control strong models with weak supervisors? A core challenge for aligning future superhuman AI systems (superalignment) is that humans will need to supervise AI systems much smarter than them. We study a simple analogy: can small models supervise large models? We show that we can use a GPT‑2‑level model to elicit most of GPT‑4’s capabilities—close to GPT‑3.5‑level performance—generalizing correctly even to hard problems where the small model failed. This opens up a new research direction that allows us to directly tackle a central challenge of aligning future superhuman models while making iterative empirical progress today.
核心贡献 · Key contributions
提出弱到强泛化作为超级对齐的新研究方向。 Proposes weak-to-strong generalization as a new research direction for superalignment.
展示 GPT-2 级别模型能激发 GPT-4 大部分能力,达到 GPT-3.5 级别性能。 Shows GPT-2-level model can elicit most of GPT-4's capabilities, achieving GPT-3.5-level performance.
证明强模型在难题上能超越弱监督者的错误进行泛化。 Demonstrates strong models generalize beyond weak supervisor's errors on hard problems.
引入简单方法鼓励强模型更自信,提升泛化能力。 Introduces a simple method encouraging strong models to be more confident, improving generalization.
提供经验证据表明朴素 RLHF 可能难以扩展到超人类模型。 Provides empirical evidence that naive RLHF may scale poorly to superhuman models.
通过当前模型作为类比,为超级对齐开辟迭代经验进展。 Opens up iterative empirical progress on superalignment by using current models as analogies.