审慎对齐:推理能力使语言模型更安全

Deliberative alignment: reasoning enables safer language models

OpenAI OpenAI · OpenAI · 2024-12-20 · OpenAI Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们为 o 系列模型引入了新的对齐策略,直接教授安全规范以及如何对其进行推理。我们提出了审慎对齐,这是一种训练范式,直接教导推理型大语言模型人类编写且可解释的安全规范文本,并训练它们在回答前显式地推理这些规范。我们使用审慎对齐来对齐 OpenAI 的 o 系列模型,使其能够利用思维链推理来反思用户提示,识别 OpenAI 内部政策中的相关文本,并起草更安全的响应。我们的方法实现了对 OpenAI 安全政策的高度精确遵守,且无需人工标注的思维链或答案。我们发现,o1 在内部和外部的多项安全基准测试中显著优于 GPT-4o 及其他最先进的大语言模型,并在许多具有挑战性的数据集上达到了饱和性能。我们相信这为提升安全性提供了一条令人兴奋的新路径,并且我们认为这是一个鼓舞人心的例子,展示了能力提升如何也能用于提升安全性。

Introducing our new alignment strategy for o-series models, which are directly taught safety specifications and how to reason over them. We introduce deliberative alignment, a training paradigm that directly teaches reasoning LLMs the text of human-written and interpretable safety specifications, and trains them to reason explicitly about these specifications before answering. We used deliberative alignment to align OpenAI’s o-series models, enabling them to use chain-of-thought (CoT) reasoning to reflect on user prompts, identify relevant text from OpenAI’s internal policies, and draft safer responses. Our approach achieves highly precise adherence to OpenAI’s safety policies, and without requiring human-labeled CoTs or answers. We find that o1 dramatically outperforms GPT‑4o and other state-of-the art LLMs across a range of internal and external safety benchmarks, and saturates performance on many challenging datasets.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 5)

阅读逐段中英对照全文 →