红队测试语言模型以减少危害:方法、扩展行为与经验教训

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Anthropic Anthropic · Anthropic · 2022-08-23 · arXiv:2209.07858 ↗ · 被引 822

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们描述了早期对语言模型进行红队测试的努力,以同时发现、衡量并尝试减少其潜在的有害输出。我们做出了三项主要贡献。首先,我们研究了 3 种模型规模(2.7B、13B 和 52B 参数)和 4 种模型类型的红队测试扩展行为:普通语言模型(LM);被提示为有帮助、诚实且无害的 LM;使用拒绝采样的 LM;以及使用人类反馈强化学习(RLHF)训练为有帮助且无害的模型。我们发现 RLHF 模型随着规模扩大越来越难以进行红队测试,而其他模型类型则呈现平稳趋势。其次,我们发布了包含 38,961 次红队攻击的数据集,供他人分析和学习。我们提供了自己的数据分析,发现了各种有害输出,从攻击性语言到更微妙的非暴力不道德输出。第三,我们详尽描述了我们的指令、流程、统计方法以及关于红队测试的不确定性。我们希望这种透明度能够加速我们作为社区共同工作的能力,以制定共享的规范、实践和技术标准,用于如何对语言模型进行红队测试。

We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main contributions. First, we investigate scaling behaviors for red teaming across 3 model sizes (2.7B, 13B, and 52B parameters) and 4 model types: a plain language model (LM); an LM prompted to be helpful, honest, and harmless; an LM with rejection sampling; and a model trained to be helpful and harmless using reinforcement learning from human feedback (RLHF). We find that the RLHF models are increasingly difficult to red team as they scale, and we find a flat trend with scale for the other model types. Second, we release our dataset of 38,961 red team attacks for others to analyze and learn from. We provide our own analysis of the data and find a variety of harmful outputs, which range from offensive language to more subtly harmful non-violent unethical outputs. Third, we exhaustively describe our instructions, processes, statistical methodologies, and uncertainty about red teaming. We hope that this transparency accelerates our ability to work together as a community in order to develop shared norms, practices, and technical standards for how to red team language models.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 19)

阅读逐段中英对照全文 →