衡量大型语言模型可扩展监督的进展

Measuring Progress on Scalable Oversight for Large Language Models

白云涛 Yuntao Bai · Anthropic · 2022-11-04 · arXiv:2211.03540 ↗ · 被引 218

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

开发安全且有用的通用人工智能系统需要我们在可扩展监督问题上取得进展,即监督可能在大多数相关技能上超越我们的系统。由于目前尚无广泛超越人类能力的系统,该问题的实证研究并不直接。本文讨论了我们思考这一问题的主要方式之一,重点关注如何通过实验进行研究。我们首先提出一个实验设计,其核心是那些人类专家能够成功但未经辅助的人类和当前通用 AI 系统失败的任务。接着,我们展示了一个概念验证实验,旨在证明该实验设计的关键特征,并通过两个问答任务(MMLU 和限时 QuALITY)展示其可行性。在这些任务中,我们发现通过聊天与不可靠的大型语言模型对话助手交互的人类参与者——这是一种可扩展监督的简单基线策略——其表现显著优于单独模型和未经辅助的人类自身。这些结果令人鼓舞,表明可扩展监督问题可以用现有模型进行研究,并支持了近期关于大型语言模型能有效协助人类完成困难任务的发现。

Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first present an experimental design centered on tasks for which human specialists succeed but unaided humans and current general AI systems fail. We then present a proof-of-concept experiment meant to demonstrate a key feature of this experimental design and show its viability with two question-answering tasks: MMLU and time-limited QuALITY. On these tasks, we find that human participants who interact with an unreliable large-language-model dialog assistant through chat -- a trivial baseline strategy for scalable oversight -- substantially outperform both the model alone and their own unaided performance. These results are an encouraging sign that scalable oversight will be tractable to study with present models and bolster recent findings that large language models can productively assist humans with difficult tasks.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 12)

阅读逐段中英对照全文 →