语言模型(大多)知道它们知道什么

Language Models (Mostly) Know What They Know

Anthropic Anthropic · Anthropic · 2022-07-11 · arXiv:2207.05221 ↗ · 被引 1729

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们研究了语言模型是否能评估自身主张的有效性,并预测哪些问题能正确回答。首先,我们发现,当以正确格式提供时,较大的模型在多样化的多项选择和真假问题上校准良好。因此,我们可以通过让模型先提出答案,然后评估其答案正确的概率“P(True)”来接近开放采样任务上的自我评估。我们发现,在多样化的任务上,P(True)在性能、校准和扩展性方面表现令人鼓舞。当允许模型在预测某一特定可能性的有效性之前考虑自己的多个样本时,自我评估的性能进一步提升。接下来,我们研究了模型是否可以被训练来预测“P(IK)”,即“我知道”问题答案的概率,而不参考任何特定提出的答案。模型在预测 P(IK)方面表现良好,并在任务间部分泛化,尽管在新任务上 P(IK)的校准存在困难。预测的 P(IK)概率在上下文中存在相关源材料时,以及在数学文字题提示解决方案时,也会适当增加。我们希望这些观察为训练更诚实的模型奠定基础,并研究当模型在模仿人类写作之外的目标上训练时,诚实如何泛化。

We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling tasks by asking models to first propose answers, and then to evaluate the probability "P(True)" that their answers are correct. We find encouraging performance, calibration, and scaling for P(True) on a diverse array of tasks. Performance at self-evaluation further improves when we allow models to consider many of their own samples before predicting the validity of one specific possibility. Next, we investigate whether models can be trained to predict "P(IK)", the probability that "I know" the answer to a question, without reference to any particular proposed answer. Models perform well at predicting P(IK) and partially generalize across tasks, though they struggle with calibration of P(IK) on new tasks. The predicted P(IK) probabilities also increase appropriately in the presence of relevant source materials in the context, and in the presence of hints towards the solution of mathematical word problems. We hope these observations lay the groundwork for training more honest models, and for investigating how honesty generalizes to cases where models are trained on objectives other than the imitation of human writing.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 38)

阅读逐段中英对照全文 →