大型语言模型的涌现能力

Emergent Abilities of Large Language Models

杰森·魏 Jason Wei · Google · 2022-06-15 · arXiv:2206.07682 ↗ · 被引 3651

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

扩展语言模型已被证明可以可预测地提高在广泛下游任务上的性能和样本效率。然而,本文讨论了一种不可预测的现象,我们称之为大型语言模型的涌现能力。我们认为,如果一种能力在较小模型中不存在,但在较大模型中出现,那么它就是涌现的。因此,涌现能力不能简单地通过外推较小模型的性能来预测。这种涌现的存在意味着进一步扩展可能会扩大语言模型的能力范围。

Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →