Six intuitions about large language models — Jason Wei
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→如今一个悬而未决的问题是,为什么大型语言模型能表现得如此出色。在这篇博客中,我将讨论关于大型语言模型的六个基本直觉。其中许多直觉源于手动检查数据,我发现这种做法很有帮助,并推荐大家尝试。语言模型通过预训练仅需预测文本语料中的下一个词,却从中学习到了惊人的知识。让我们来看一些例子,了解它们从这一下一词预测任务中可能学到什么。直觉一:在大型自监督数据上进行下一词预测,本质上是一种大规模的多任务学习。
An open question these days is why large language models work so well. In this blog post I will discuss six basic intuitions about large language models. Many of them are inspired by manually examining data, which is an exercise that I’ve found helpful and would recommend. Language models are pre-trained to simply predict the next word in a corpus of text, and they learn a surprising amount from this. Let’s look at some examples of what they might learn from this next-word prediction task. Intuition 1. Next-word prediction on large, self-supervised data is massively multi-task learning.
当前一个悬而未决的问题是,为什么大型语言模型表现得如此出色。在这篇博客文章中,我将讨论关于大型语言模型的六个基本直觉。其中许多直觉源于手动检查数据,我发现这种做法很有帮助,并愿意推荐。
An open question these days is why large language models work so well. In this blog post I will discuss six basic intuitions about large language models. Many of them are inspired by manually examining data, which is an exercise that I’ve found helpful and would recommend.
语言模型通过简单地预测文本语料库中的下一个词进行预训练,并从中学习到惊人的知识。让我们看一些示例,了解它们可能从这一下一个词预测任务中学到什么。
Language models are pre-trained to simply predict the next word in a corpus of text, and they learn a surprising amount from this. Let’s look at some examples of what they might learn from this next-word prediction task.
直觉一:在大规模自监督数据上进行下一个词预测本质上是大规模多任务学习。
Intuition 1. Next-word prediction on large, self-supervised data is massively multi-task learning.
尽管下一个词预测是一项极其简单的任务,但与大规模数据集结合后,它迫使模型学习大量任务。请考虑以下传统自然语言处理(NLP)任务的示例,这些任务可以通过预测语料库中某些文本的下一个词来学习。
Although next-word prediction is an extremely simple task, when combined with massive datasets, it forces the model to learn a lot of tasks. Consider the following examples of traditional NLP tasks that can be learned by predicting the next word on some text in the corpus.
上述任务明确但有点理想化。实际上,预测下一个词涉及执行许多“奇特”的任务。请考虑以下句子:
The above tasks are clear-cut but a bit idealized. In reality, predicting the next word involves doing many “odd” tasks. Consider the following sentence:
当你以这种方式看待数据时,很明显,下一个词预测迫使模型学习大量关于语言的知识;不仅是句法和语义,还包括逗号预测、事实知识,甚至可能是推理。这是一个有趣的例子,说明了一个简单的目标与复杂的数据相结合时,如何能够产生高度智能的行为(前提是你同意语言模型具有智能)。
When you view the data in this way, it is obvious that next-word prediction forces the model to learn a lot about language; not just syntax and semantics, but also things like comma prediction, factual knowledge, perhaps even reasoning. This is an interesting example of how a simple objective, when combined with complex data, can lead to highly intelligent behavior (assuming you agree that language models are intelligent).
直觉 2:学习输入-输出关系可以转化为下一个词预测。这被称为上下文学习。
Intuition 2. Learning input-output relationships can be cast as next-word prediction. This is known as in-context learning.
过去几十年的机器学习一直专注于学习 <输入, 输出> 对之间的关系。由于下一个词预测具有如此强的通用性,我们可以很容易地将机器学习转化为下一个词预测。我们称之为上下文学习(又称少样本学习或少样本提示)。这由 GPT-3 论文首创,该论文提出使用自然语言指令后跟 <输入, 输出> 对。下图左侧来自 GPT-3 论文。
The past decades of machine learning have focused on learning the relationships between <input, output> pairs. Because next-word prediction is so general, we can easily cast machine learning as next-word prediction. We call this in-context learning (a.k.a. few-shot learning or few-shot prompting). This was pioneered by the GPT-3 paper, which proposed using a natural language instruction followed by <input, output> pairs. This is shown in the left image below from the GPT-3 paper.
在上图右侧,我们可以看到,增加上下文中的示例数量会提高 GPT-3 论文中一个任务的性能。这意味着模型能从看到这些 <输入, 输出> 示例中受益。
In the right part of the image above, we can see that increasing the number of examples in context improves performance for a task in the GPT-3 paper. This means that the model benefits from seeing these <input, output> examples.
上下文学习是使用大语言模型的一种标准方式,其方便之处在于,<输入, 输出> 对正是过去几十年我们做机器学习的方式。然而,从第一性原理来看,我们并没有理由继续沿用 <输入, 输出> 对。当我们与人类交流时,我们会给他们指示、解释,并互动式地教导他们。
In-context learning is a standard formulation of using large language models that is convenient because <input, output> pairs were how we did machine learning for the past decades. However, there is no first-principles reason why we continue following <input, output> pairs. When we communicate with humans, we give them instructions, explanations, and teach them interactively.
直觉 3:词元(token)的信息密度差异很大,所以要让语言模型有时间思考。
Intuition 3. Tokens can have very different information density, so give language models time to think.
一个基本事实是,并非所有词元在信息含量上都具有同等价值。
It is a fundamental truth that not all tokens are worth the same in terms of information.
1. 有些词元非常容易猜到,几乎没什么价值。例如,在“I’m Jason Wei, a researcher at OpenAI working on large language ___”中,预测出“models”并不难。这个词元太容易预测,以至于即使省略它,也不会损失多少信息。
1. Some tokens are very easy to guess and not worth much at all. For example, in “I’m Jason Wei, a researcher at OpenAI working on large language ___”, it’s not so hard to predict “models”. It’s so easy to predict that token that not much information is lost if I omit it.
2. 有些词元非常难猜,它们价值很高。例如,在“Jason Wei’s favorite color is ___”中,这个词元几乎不可能预测。因此,这个词元包含大量新信息。
2. Some tokens are very hard to guess; they’re worth a lot. For example, in “Jason Wei’s favorite color is ___”, it’s virtually impossible to predict. So that token contains a lot of new information.
3. 有些词元也可能非常难以计算。例如,在“问题:\( \left( \frac{((8-2) \times 3+4)^3}{8} \right)^2 \) 等于多少?(A)1,483,492;(B)1,395,394;(C)1,771,561;答案:(”中,下一个词元需要大量运算(也就是计算该表达式的值)。
3. Some tokens can also be very hard to compute. For example, in “Question: What is \( \left( \frac{((8-2) \times 3+4)^3}{8} \right)^2 \)? (A) 1,483,492; (B) 1,395,394; (C) 1,771,561; Answer: (”, the next token requires a lot of work (evaluating that expression).
你可以想象,如果你是 ChatGPT,一旦看到提示词就必须立即开始打字,那么你很难正确回答那个问题。
You can imagine that if you’re ChatGPT, and as soon as you have to see the prompt you have to immediately start typing, it would be pretty hard to get that question right.
解决办法是给语言模型更多算力,允许它们在给出最终答案之前进行自然语言推理。这可以通过一个简单的技巧实现,即思维链提示(chain-of-thought prompting),它通过在小样本示例中提供一个“思维链”示例来鼓励模型进行推理,如蓝色突出显示的部分所示。
The solution to this is to give language models more compute by allowing them to perform natural language reasoning before giving the final answer. This can be done via a simple trick called chain-of-thought prompting, which encourages the model to reason by providing an example of a “chain-of-thought” in the few-shot example, as highlighted in blue.
这项技术对于提高复杂推理任务的性能非常有用,这些任务需要人类花超过一秒钟才能解决。对于比上面简单算术问题更复杂的问题,可以让语言模型先将提示词分解为子问题,然后依次解决这些子问题(“从易到难提示”,least-to-most prompting)。这种范式之所以强大,是因为我们希望 AI 最终能够解决人类面临的最困难的问题(例如贫困、气候变化等),而推理能力是解决这类问题的基本基石。
This technique is really useful for improving performance on complicated reasoning tasks that would require humans to spend more than one second solving. For even-more-complicated problems than the simple arithmetic problem shown above, it can help to have the language model decompose the prompt first into subproblems, and then sequentially solve the subproblems (“least-to-most prompting”). This paradigm is powerful because we want AI to eventually be able to solve the hardest problems we face as humans (e.g., poverty, climate change, etc), and being able to reason is a fundamental building block for solving such problems.
上述下一个词预测任务之所以有效,关键在于缩放(scaling),即用更多数据训练更大的神经网络。显然,训练前沿语言模型成本很高,而我们之所以这样做,是因为我们有合理信心:使用更大的神经网络和更多数据确实会带来更好的模型(即,当模型和数据规模增大时,性能可能不会饱和)。
The key reason that the above next-word prediction tasks work is scaling, which means training larger neural nets on more data. Obviously it costs a lot of money to train frontier language models, and so the reason we do it is that we have reasonable confidence that using larger neural networks and more data will actually lead to a better model (i.e., performance probably won’t saturate when you increase the model and data size).
直觉 4:缩放语言模型(规模和数据)预计将持续改善损失。
Intuition 4. Scaling language models (size and data) is expected to continue improving loss.
Scaling(规模扩张)能可预测地提升性能,这一现象被称为“缩放定律”;如下左图所示,随着算力增加,测试损失平滑下降。
The fact that scaling improves performance predictably is called “scaling laws,” as shown in the left figure below: as you increase compute, test loss improves smoothly.
右图是另一个证据,表明随着语言模型规模扩展,损失如何平滑下降——通过追踪较小模型的损失曲线,你最多只需 1/10,000 的算力就能预测 GPT-4 的损失。
The right figure is another piece of evidence of how loss smoothly improves as you scale up the language model—by tracing the loss curve of smaller models, you can predict GPT-4’s loss using up to 10,000× less compute.
Scaling(规模扩张)究竟为何有效仍是一个开放问题,但这里有两条不太严谨的解释。其一,小语言模型无法在其参数中记住那么多知识,而大语言模型可以记住关于世界的大量事实信息。其二,小语言模型受容量限制,可能只能学习数据中的一阶相关性;相比之下,大语言模型可以学习数据中的复杂启发式规则。
It is an open question why exactly scaling works, but here are two hand-wavy reasons. One is that small language models can’t memorize as much knowledge in their parameters, whereas large language models can memorize a huge amount of factual information about the world. A second guess is that while small language models are capacity-constrained, they might only learn first-order correlations in data. Large language models, on the other hand, can learn complex heuristics in data.
直觉 5:虽然整体损失随规模扩展平滑下降,但单个下游任务可能以涌现的方式变化。
Intuition 5. While overall loss scales smoothly, individual downstream tasks may scale in an emergent fashion.
让我们仔细看看损失下降时究竟会发生什么。你可以将整体损失视为大量已学习任务的加权平均,例如,
Let’s take a closer look at what exactly happens when loss improves. You can consider overall loss as a weighted average of the massive amount of tasks learned, e.g.,
\(L_{\text{总体损失}} = 10^{-10} \cdot L_{\text{语法任务}} + 10^{-10} \cdot L_{\text{情感分析任务}} + \dots\)
\(L_{\text{overall}} = 10^{-10} \cdot L_{\text{grammar}} + 10^{-10} \cdot L_{\text{sentiment}} + \dots\)
\+ 10^{-10} \cdot L_{\text{数学能力任务}} + \dots\
\+ 10^{-10} \cdot L_{\text{math}} + \dots\
现在考虑你的损失从 4 降到 3。所有任务是否都均匀地变好了?很可能不是。也许损失 = 4 的模型的语法已经完美,因此该项已经饱和,但损失 = 3 的模型在数学能力上提升了很多。
Now consider your loss going from 4 to 3. Do all tasks get better uniformly? Probably not. Maybe the grammar of the model with loss = 4 was already perfect, so that is saturated, but the math ability improves a lot in the model with loss = 3.
事实证明,如果你观察模型在 200 个下游任务上的表现,你会发现有些任务平滑地改善,另一些任务完全没有改善,还有一些任务突然改善。这里就是这类任务的八个例子:小模型的性能大约等于随机水平,而一旦模型规模达到某个阈值,性能就会大幅超过随机水平。
It turns out that if you look at the performance of the model on 200 downstream tasks, you’ll see that while some tasks smoothly improve, other tasks don’t improve at all, and some tasks improve suddenly. Here are eight examples of such tasks, where performance is about random for small models, and increases to well above random once the model size reaches a certain threshold.
我们用“涌现”来指代由量变引发的质变。更具体地说,如果一个能力在较小模型中不存在,而在较大模型中存在,我们就称该大型语言模型的能力具有涌现性。在这类任务中,我们常常看到:较小模型的性能大约等于随机水平,而大于某个阈值规模的模型的性能则显著高于随机水平,如下图所示。
The term we use for qualitative changes arising from quantitative changes is “emergence”. More specifically, we call an ability of a large language model emergent if it is not present in smaller models, but is present in larger models. In such tasks, we often see that performance is about random for small models and well above random for models larger than a certain threshold size, as shown in the figure below.
涌现有三个重要含义:
There are three important implications of emergence:
1. 涌现不能通过简单地从较小模型外推缩放曲线来预测。
1. Emergence is not predictable by simply extrapolating scaling curves from smaller models.
2. 涌现能力并非由语言模型的训练者显式指定。
2. Emergent abilities are not explicitly specified by the trainer of the language model.
3. 既然规模扩张已经解锁了涌现能力,可以预期进一步的规模扩张会引发更多能力。
3. Since scaling has unlocked emergent abilities, further scaling can be expected to further elicit more abilities.
直觉 6:真正的上下文学习会发生,但仅在足够大的语言模型中。
_Intuition 6. Real in-context learning happens, but only in large-enough language models._
我们从 GPT-3 论文中看到,增加上下文示例的数量可以提升性能。虽然我们希望这是因为模型确实从上下文中的示例学习了 <输入, 输出> 映射,但性能提升也可能源于其他原因,例如示例向模型传达了格式或可能的标签信息。
We have seen from the GPT-3 paper that increasing the number of in-context examples improves performance. While we hope that this is because the model actually learns <input, output> mappings from the examples in its context, the improvement in performance could be due to other reasons, such as the examples telling the model about formatting or possible labels.
事实上,有论文表明,即使对上下文示例使用随机标签,GPT-3 的性能也几乎没有下降。作者认为,性能提升因此并非来自学习 <输入, 输出> 映射,而是来自上下文示例所传授的格式或可能标签等信息。
In fact, one paper showed that GPT-3’s performance barely decreases even if you use random labels for the in-context examples. The authors suggest that the performance improvement is therefore not due to learning <input, output> mappings, but rather due to the in-context examples teaching things like formatting or the possible labels.
然而,与当今最强大的模型相比,GPT-3 并不算是一个超级“大”语言模型。如果我们采用更极端的标签翻转设置(即正例表示负例,负例表示正例),就会发现大语言模型强烈遵循翻转后的标签,而小语言模型的性能则完全不受影响。如下图所示,大语言模型(PaLM-540B、code-davinci-002 和 text-davinci-002)的性能出现了下降。
However, GPT-3 is not a super “large” language model compared to the most powerful models today. If we take a more extreme setting of flipped labels (i.e., positive means negative and negative means positive), then we find that large language models strongly follow the flipped labels, while small language models are not affected in performance at all. This is shown in the figure below, where performance dips for the large language models (PaLM-540B, code-davinci-002, and text-davinci-002).
这里的要点是,语言模型确实会关注 <输入, 输出> 映射,但前提是语言模型足够大。
The takeaway here is that language models do look at <input, output> mappings, but only if the language model is large enough.
我希望上述直觉尽管基础,但仍然有用。贯穿许多直觉的一个共同主题是,通过手动查看数据,你可以学到很多东西。我最近很喜欢这样做,并强烈推荐 :)
I hope the above intuitions were useful despite how basic they are. One theme that is common across many of the intuitions is that you can learn a lot by manually looking at data. I enjoyed doing this recently and highly recommend it :)
上一篇:成功的语言模型评估 | 下一篇:对 Twitter 追踪的观察
Previous: Successful language model evals | Next: Observations from tracking Twitter