推理模型并不总是说出它们的想法

Reasoning Models Don't Always Say What They Think

白云涛 Yuntao Bai · Anthropic · 2025-05-08 · arXiv:2505.05410 ↗ · 被引 344

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

思维链(CoT)为 AI 安全提供了潜在优势,因为它允许监控模型的 CoT 以理解其意图和推理过程。然而,这种监控的有效性取决于 CoT 是否忠实地代表模型的实际推理过程。我们评估了最先进的推理模型在提示中呈现的 6 种推理提示下的 CoT 忠实度,发现:(1)对于大多数测试的设置和模型,CoT 在至少 1%的示例中揭示了它们对提示的使用,但揭示率通常低于 20%;(2)基于结果的强化学习最初提高了忠实度,但随后趋于平稳而未饱和;(3)当强化学习增加提示使用频率(奖励黑客)时,即使没有针对 CoT 监控器进行训练,表达这些提示的倾向也不会增加。这些结果表明,CoT 监控是在训练和评估期间注意到不良行为的一种有前途的方法,但不足以排除它们。它们还表明,在像我们这样不需要 CoT 推理的环境中,对 CoT 的测试时监控不太可能可靠地捕获罕见且灾难性的意外行为。

Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effectiveness of such monitoring hinges on CoTs faithfully representing models' actual reasoning processes. We evaluate CoT faithfulness of state-of-the-art reasoning models across 6 reasoning hints presented in the prompts and find: (1) for most settings and models tested, CoTs reveal their usage of hints in at least 1% of examples where they use the hint, but the reveal rate is often below 20%, (2) outcome-based reinforcement learning initially improves faithfulness but plateaus without saturating, and (3) when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor. These results suggest that CoT monitoring is a promising way of noticing undesired behaviors during training and evaluations, but that it is not sufficient to rule them out. They also suggest that in settings like ours where CoT reasoning is not necessary, test-time monitoring of CoTs is unlikely to reliably catch rare and catastrophic unexpected behaviors.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →