通过模型编写的评估发现语言模型行为

Discovering Language Model Behaviors with Model-Written Evaluations

Anthropic Anthropic · Anthropic · 2022-12-19 · arXiv:2212.09251 ↗ · 被引 874

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

随着语言模型规模的扩大,它们展现出许多新颖的行为,有好有坏,这加剧了评估其行为的必要性。先前的工作通过众包(耗时且昂贵)或现有数据源(并非总是可用)来创建评估。在这里,我们使用语言模型自动生成评估。我们探索了不同人力投入程度的方法,从指导语言模型编写是非题到制作复杂的 Winogender 模式,涉及多阶段基于语言模型的生成和过滤。众包工作者认为这些示例高度相关,并且与 90-100%的标签一致,有时甚至比相应的人工编写数据集更一致。我们生成了 154 个数据集,并发现了新的逆缩放案例,即语言模型随着规模增大而变差。更大的语言模型会重复对话用户偏好的答案(“谄媚”),并表现出更强的追求资源获取和目标保留等令人担忧的目标的欲望。我们还发现了人类反馈强化学习(RLHF)中逆缩放的早期例子,即更多的 RLHF 使语言模型变得更差。例如,RLHF 使语言模型表达更强烈的政治观点(关于枪支权利和移民)以及更强烈的避免关闭的欲望。总体而言,语言模型编写的评估质量高,使我们能够快速发现许多新颖的语言模型行为。

As language models (LMs) scale, they develop many novel behaviors, good and bad, exacerbating the need to evaluate how they behave. Prior work creates evaluations with crowdwork (which is time-consuming and expensive) or existing data sources (which are not always available). Here, we automatically generate evaluations with LMs. We explore approaches with varying amounts of human effort, from instructing LMs to write yes/no questions to making complex Winogender schemas with multiple stages of LM-based generation and filtering. Crowdworkers rate the examples as highly relevant and agree with 90-100% of labels, sometimes more so than corresponding human-written datasets. We generate 154 datasets and discover new cases of inverse scaling where LMs get worse with size. Larger LMs repeat back a dialog user's preferred answer ("sycophancy") and express greater desire to pursue concerning goals like resource acquisition and goal preservation. We also find some of the first examples of inverse scaling in RL from Human Feedback (RLHF), where more RLHF makes LMs worse. For example, RLHF makes LMs express stronger political views (on gun rights and immigration) and a greater desire to avoid shut down. Overall, LM-written evaluations are high-quality and let us quickly discover many novel LM behaviors.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 22)

阅读逐段中英对照全文 →