通用语言助手作为对齐研究的实验室

A General Language Assistant as a Laboratory for Alignment

白云涛 Yuntao Bai · Anthropic · 2021-12-01 · arXiv:2112.00861 ↗ · 被引 1155

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

鉴于大型语言模型的广泛能力,我们应致力于开发一种通用、基于文本的助手,使其与人类价值观对齐,即有用、诚实且无害。作为这一方向的初步探索,我们研究了简单的基线技术和评估方法,如提示。我们发现,适度的干预措施带来的益处随模型规模增大而增加,能泛化到多种对齐评估中,且不会损害大型模型的性能。接下来,我们研究了与对齐相关的几种训练目标的扩展趋势,比较了模仿学习、二元判别和排序偏好建模。我们发现,排序偏好建模的表现远优于模仿学习,且通常随模型规模扩展更有利。相比之下,二元判别的表现和扩展趋势与模仿学习非常相似。最后,我们研究了一个“偏好模型预训练”阶段,旨在提高在人类偏好上微调时的样本效率。

Given the broad capabilities of large language models, it should be possible to work towards a general-purpose, text-based assistant that is aligned with human values, meaning that it is helpful, honest, and harmless. As an initial foray in this direction we study simple baseline techniques and evaluations, such as prompting. We find that the benefits from modest interventions increase with model size, generalize to a variety of alignment evaluations, and do not compromise the performance of large models. Next we investigate scaling trends for several training objectives relevant to alignment, comparing imitation learning, binary discrimination, and ranked preference modeling. We find that ranked preference modeling performs much better than imitation learning, and often scales more favorably with model size. In contrast, binary discrimination typically performs and scales very similarly to imitation learning. Finally we study a `preference model pre-training' stage of training, with the goal of improving sample efficiency when finetuning on human preferences.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 40)

阅读逐段中英对照全文 →