弱到强泛化

Weak-to-strong generalization

OpenAI OpenAI · OpenAI · 2023-12-14 · OpenAI Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了一个用于超级对齐的新研究方向,并附有初步的有希望的结果:我们能否利用深度学习的泛化特性,用弱监督者控制强模型?对齐未来超人类 AI 系统(超级对齐)的一个核心挑战是,人类需要监督比他们聪明得多的 AI 系统。我们研究了一个简单的类比:小模型能否监督大模型?我们表明,我们可以使用 GPT-2 级别的模型来引出 GPT-4 的大部分能力——接近 GPT-3.5 级别的性能——即使在小模型失败的难题上也能正确泛化。这开启了一个新的研究方向,使我们能够直接应对对齐未来超人类模型的核心挑战,同时今天就能进行迭代的实证进展。

We present a new research direction for superalignment, together with promising initial results: can we leverage the generalization properties of deep learning to control strong models with weak supervisors? A core challenge for aligning future superhuman AI systems (superalignment) is that humans will need to supervise AI systems much smarter than them. We study a simple analogy: can small models supervise large models? We show that we can use a GPT‑2‑level model to elicit most of GPT‑4’s capabilities—close to GPT‑3.5‑level performance—generalizing correctly even to hard problems where the small model failed. This opens up a new research direction that allows us to directly tackle a central challenge of aligning future superhuman models while making iterative empirical progress today.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 5)

阅读逐段中英对照全文 →