理解 LSTM 网络

Understanding LSTM Networks

克里斯·奥拉 Chris Olah · Anthropic · 2015-08-27 · Colah's Blog ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

人类不会每秒钟从头开始思考。当你阅读这篇文章时,你基于对前面单词的理解来理解每个单词。你不会抛弃一切重新开始思考。你的思想具有持久性。传统的神经网络无法做到这一点,这似乎是一个重大缺陷。例如,想象一下你想对电影中每个时刻发生的事件进行分类。不清楚传统神经网络如何利用其对电影中先前事件的推理来影响后续事件。循环神经网络解决了这个问题。它们是带有循环的网络,允许信息持久化。

Humans don’t start their thinking from scratch every second. As you read this essay, you understand each word based on your understanding of previous words. You don’t throw everything away and start thinking from scratch again. Your thoughts have persistence. Traditional neural networks can’t do this, and it seems like a major shortcoming. For example, imagine you want to classify what kind of event is happening at every point in a movie. It’s unclear how a traditional neural network could use its reasoning about previous events in the film to inform later ones. Recurrent neural networks address this issue. They are networks with loops in them, allowing information to persist.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 7)

全文 · Full text(逐段中英对照)

循环神经网络 Recurrent Neural Networks

人类不会每秒钟都从头开始思考。当你阅读这篇文章时,你会基于对前面词语的理解来理解每个词。你不会抛弃一切,从头开始思考。你的思维具有持久性。

Humans don’t start their thinking from scratch every second. As you read this essay, you understand each word based on your understanding of previous words. You don’t throw everything away and start thinking from scratch again. Your thoughts have persistence.

传统神经网络无法做到这一点,这似乎是一个重大缺陷。例如,假设你想对电影中每个时刻发生的事件进行分类。传统神经网络如何利用其对电影中先前事件的推理来影响后续事件,这一点尚不清楚。

Traditional neural networks can’t do this, and it seems like a major shortcoming. For example, imagine you want to classify what kind of event is happening at every point in a movie. It’s unclear how a traditional neural network could use its reasoning about previous events in the film to inform later ones.

循环神经网络解决了这个问题。它们是带有循环的网络,允许信息持续存在。

Recurrent neural networks address this issue. They are networks with loops in them, allowing information to persist.

在上图中,一个神经网络块 A$A$ 查看某个输入 x t$x_{t}$ 并输出一个值 h t$h_{t}$。循环允许信息从网络的一个步骤传递到下一个步骤。

In the above diagram, a chunk of neural network, A$A$, looks at some input x t$x_{t}$ and outputs a value h t$h_{t}$. A loop allows information to be passed from one step of the network to the next.

这些循环使得循环神经网络看起来有些神秘。然而,如果你再仔细想想,就会发现它们与普通神经网络并没有太大不同。循环神经网络可以被视为同一网络的多个副本,每个副本向后继者传递消息。考虑一下如果我们展开循环会发生什么:

These loops make recurrent neural networks seem kind of mysterious. However, if you think a bit more, it turns out that they aren’t all that different than a normal neural network. A recurrent neural network can be thought of as multiple copies of the same network, each passing a message to a successor. Consider what happens if we unroll the loop:

这种链式性质揭示了循环神经网络与序列和列表密切相关。它们是用于此类数据的神经网络的自然架构。

This chain-like nature reveals that recurrent neural networks are intimately related to sequences and lists. They’re the natural architecture of neural network to use for such data.

而且它们确实被使用了!在过去的几年里,将 RNN 应用于各种问题取得了令人难以置信的成功:语音识别、语言建模、翻译、图像描述……不胜枚举。我将把使用 RNN 可以实现的惊人成就留给 Andrej Karpathy 的优秀博客文章《循环神经网络的非凡有效性》来讨论。但它们确实非常了不起。

And they certainly are used! In the last few years, there have been incredible success applying RNNs to a variety of problems: speech recognition, language modeling, translation, image captioning… The list goes on. I’ll leave discussion of the amazing feats one can achieve with RNNs to Andrej Karpathy’s excellent blog post, The Unreasonable Effectiveness of Recurrent Neural Networks. But they really are pretty amazing.

这些成功的关键在于使用了“LSTM”,这是一种非常特殊的循环神经网络,在许多任务上比标准版本好得多。几乎所有基于循环神经网络的激动人心的结果都是通过它们实现的。本文将要探讨的正是这些 LSTM。

Essential to these successes is the use of “LSTMs,” a very special kind of recurrent neural network which works, for many tasks, much much better than the standard version. Almost all exciting results based on recurrent neural networks are achieved with them. It’s these LSTMs that this essay will explore.

长期依赖问题 The Problem of Long-Term Dependencies

RNN 的吸引力之一在于,它们或许能够将先前信息与当前任务联系起来,例如利用之前的视频帧来帮助理解当前帧。如果 RNN 能做到这一点,它们将非常有用。但它们能做到吗?这取决于具体情况。

One of the appeals of RNNs is the idea that they might be able to connect previous information to the present task, such as using previous video frames might inform the understanding of the present frame. If RNNs could do this, they’d be extremely useful. But can they? It depends.

有时,我们只需要查看最近的信息就能完成当前任务。例如,考虑一个语言模型试图根据前面的词预测下一个词。如果我们试图预测“the clouds are in the _sky_”中的最后一个词,我们不需要任何进一步的上下文——很明显下一个词是“sky”。在这种情况下,相关信息与所需位置之间的间隔很小,RNN 可以学会利用过去的信息。

Sometimes, we only need to look at recent information to perform the present task. For example, consider a language model trying to predict the next word based on the previous ones. If we are trying to predict the last word in “the clouds are in the _sky_,” we don’t need any further context – it’s pretty obvious the next word is going to be sky. In such cases, where the gap between the relevant information and the place that it’s needed is small, RNNs can learn to use the past information.

但也存在需要更多上下文的情况。考虑预测文本“I grew up in France… I speak fluent _French_”中的最后一个词。最近的信息表明下一个词很可能是一种语言的名称,但如果我们想缩小范围确定是哪种语言,我们需要更早的“France”这一上下文。相关信息与所需位置之间的间隔完全可能变得非常大。

But there are also cases where we need more context. Consider trying to predict the last word in the text “I grew up in France… I speak fluent _French_.” Recent information suggests that the next word is probably the name of a language, but if we want to narrow down which language, we need the context of France, from further back. It’s entirely possible for the gap between the relevant information and the point where it is needed to become very large.

不幸的是,随着间隔的增大,RNN 无法学会连接这些信息。

Unfortunately, as that gap grows, RNNs become unable to learn to connect the information.

理论上,RNN 完全能够处理这种“长期依赖”。人类可以仔细选择参数,让它们解决这种形式的玩具问题。但遗憾的是,在实践中,RNN 似乎无法学会它们。这个问题由 Hochreiter (1991) [德文] 和 Bengio 等人 (1994) 深入探讨,他们发现了一些根本性的原因,解释了为什么这可能很困难。

In theory, RNNs are absolutely capable of handling such “long-term dependencies.” A human could carefully pick parameters for them to solve toy problems of this form. Sadly, in practice, RNNs don’t seem to be able to learn them. The problem was explored in depth by [Hochreiter (1991) [German]](http://people.idsia.ch/~juergen/SeppHochreiter1991ThesisAdvisorSchmidhuber.pdf) and Bengio, et al. (1994), who found some pretty fundamental reasons why it might be difficult.

幸运的是,LSTM 没有这个问题!

Thankfully, LSTMs don’t have this problem!

LSTM 网络 LSTM Networks

长短期记忆网络——通常简称为“LSTM”——是一种特殊的循环神经网络(RNN),能够学习长期依赖关系。它由 Hochreiter 和 Schmidhuber(1997)提出,并在后续工作中被许多人改进和推广。1 它们在大量不同的问题上表现极佳,现已广泛使用。

Long Short Term Memory networks – usually just called “LSTMs” – are a special kind of RNN, capable of learning long-term dependencies. They were introduced by Hochreiter & Schmidhuber (1997), and were refined and popularized by many people in following work.1 They work tremendously well on a large variety of problems, and are now widely used.

LSTM 被明确设计用来避免长期依赖问题。长时间记住信息实际上是它们的默认行为,而不是它们难以学习的事情!

LSTMs are explicitly designed to avoid the long-term dependency problem. Remembering information for long periods of time is practically their default behavior, not something they struggle to learn!

所有循环神经网络都具有神经网络重复模块链的形式。在标准 RNN 中,这个重复模块具有非常简单的结构,例如单个 tanh 层。

All recurrent neural networks have the form of a chain of repeating modules of neural network. In standard RNNs, this repeating module will have a very simple structure, such as a single tanh layer.

标准 RNN 中的重复模块包含单个层。

The repeating module in a standard RNN contains a single layer.

LSTM 也具有这种链式结构,但重复模块具有不同的结构。它不是一个单一的神经网络层,而是四个以特殊方式相互作用的层。

LSTMs also have this chain like structure, but the repeating module has a different structure. Instead of having a single neural network layer, there are four, interacting in a very special way.

LSTM 中的重复模块包含四个相互作用的层。

The repeating module in an LSTM contains four interacting layers.

不必担心具体细节。我们稍后将逐步讲解 LSTM 图。现在,让我们先熟悉一下我们将使用的符号。

Don’t worry about the details of what’s going on. We’ll walk through the LSTM diagram step by step later. For now, let’s just try to get comfortable with the notation we’ll be using.

在上图中,每条线携带一个完整的向量,从一个节点的输出到其他节点的输入。粉色圆圈表示逐点操作,如向量加法,而黄色框是学习的神经网络层。线合并表示连接,而线分叉表示其内容被复制,副本被发送到不同位置。

In the above diagram, each line carries an entire vector, from the output of one node to the inputs of others. The pink circles represent pointwise operations, like vector addition, while the yellow boxes are learned neural network layers. Lines merging denote concatenation, while a line forking denote its content being copied and the copies going to different locations.

LSTM 背后的核心思想 The Core Idea Behind LSTMs

LSTM 的关键是细胞状态,即图中贯穿顶部的水平线。

The key to LSTMs is the cell state, the horizontal line running through the top of the diagram.

细胞状态有点像传送带。它直接沿着整个链条运行,只有一些轻微的线性交互。信息很容易沿着它不变地流动。

The cell state is kind of like a conveyor belt. It runs straight down the entire chain, with only some minor linear interactions. It’s very easy for information to just flow along it unchanged.

LSTM 确实有能力移除或添加信息到细胞状态,这种能力由称为门的结构精心调节。

The LSTM does have the ability to remove or add information to the cell state, carefully regulated by structures called gates.

门是一种可选地让信息通过的方式。它们由一个 sigmoid 神经网络层和一个逐点乘法操作组成。

Gates are a way to optionally let information through. They are composed out of a sigmoid neural net layer and a pointwise multiplication operation.

sigmoid 层输出介于 0 和 1 之间的数字,描述每个组件应该被允许通过多少。值为 0 表示“什么都不让通过”,而值为 1 表示“让一切通过!”

The sigmoid layer outputs numbers between zero and one, describing how much of each component should be let through. A value of zero means “let nothing through,” while a value of one means “let everything through!”

LSTM 有三个这样的门,用于保护和控制细胞状态。

An LSTM has three of these gates, to protect and control the cell state.

逐步 LSTM 详解 Step-by-Step LSTM Walk Through

LSTM 的第一步是决定我们要从细胞状态中丢弃哪些信息。这个决定由一个称为“遗忘门层”的 sigmoid 层做出。它查看 h_{t-1}和 x_t,并为细胞状态 C_{t-1}中的每个数字输出一个介于 0 和 1 之间的数。1 表示“完全保留”,0 表示“完全丢弃”。

The first step in our LSTM is to decide what information we’re going to throw away from the cell state. This decision is made by a sigmoid layer called the “forget gate layer.” It looks at h t−1$h_{t - 1}$ and x t$x_{t}$, and outputs a number between 0$0$ and 1$1$ for each number in the cell state C t−1$C_{t - 1}$. A 1$1$ represents “completely keep this” while a 0$0$ represents “completely get rid of this.”

让我们回到语言模型根据之前所有词预测下一个词的例子。在这样的问题中,细胞状态可能包含当前主语的性质,以便使用正确的代词。当我们看到一个新的主语时,我们希望忘记旧主语的性质。

Let’s go back to our example of a language model trying to predict the next word based on all the previous ones. In such a problem, the cell state might include the gender of the present subject, so that the correct pronouns can be used. When we see a new subject, we want to forget the gender of the old subject.

下一步是决定我们要在细胞状态中存储哪些新信息。这包括两部分。首先,一个称为“输入门层”的 sigmoid 层决定我们要更新哪些值。接着,一个 tanh 层创建一个新的候选值向量\tilde{C}_t,它可以被添加到状态中。在下一步中,我们将结合这两者来创建对状态的更新。

The next step is to decide what new information we’re going to store in the cell state. This has two parts. First, a sigmoid layer called the “input gate layer” decides which values we’ll update. Next, a tanh layer creates a vector of new candidate values, C~t$\overset{\sim}{C}_{t}$, that could be added to the state. In the next step, we’ll combine these two to create an update to the state.

在我们的语言模型例子中,我们希望将新主语的性质添加到细胞状态中,以替换我们正在忘记的旧主语的性质。

In the example of our language model, we’d want to add the gender of the new subject to the cell state, to replace the old one we’re forgetting.

现在是时候将旧的细胞状态 C_{t-1}更新为新的细胞状态 C_t 了。前面的步骤已经决定了要做什么,我们只需要实际执行它。

It’s now time to update the old cell state, C t−1$C_{t - 1}$, into the new cell state C t$C_{t}$. The previous steps already decided what to do, we just need to actually do it.

我们将旧状态乘以 f_t,忘记我们之前决定要忘记的东西。然后我们加上 i_t * \tilde{C}_t。这是新的候选值,按我们决定更新每个状态值的程度进行缩放。

We multiply the old state by f t$f_{t}$, forgetting the things we decided to forget earlier. Then we add i t∗C~t$i_{t} * \overset{\sim}{C}_{t}$. This is the new candidate values, scaled by how much we decided to update each state value.

在语言模型的例子中,这正是我们实际丢弃关于旧主语性质的信息并添加新信息的地方,正如我们在前面步骤中决定的那样。

In the case of the language model, this is where we’d actually drop the information about the old subject’s gender and add the new information, as we decided in the previous steps.

最后,我们需要决定我们要输出什么。这个输出将基于我们的细胞状态,但会是一个过滤后的版本。首先,我们运行一个 sigmoid 层,它决定我们要输出细胞状态的哪些部分。然后,我们将细胞状态通过 tanh(将值压缩到-1 和 1 之间),并乘以 sigmoid 门的输出,这样我们只输出我们决定的部分。

Finally, we need to decide what we’re going to output. This output will be based on our cell state, but will be a filtered version. First, we run a sigmoid layer which decides what parts of the cell state we’re going to output. Then, we put the cell state through tanh$tanh$ (to push the values to be between −1$- 1$ and 1$1$) and multiply it by the output of the sigmoid gate, so that we only output the parts we decided to.

对于语言模型例子,由于它刚刚看到了一个主语,它可能想要输出与动词相关的信息,以防接下来出现动词。例如,它可能输出主语是单数还是复数,这样如果接下来是动词,我们就知道动词应该变为什么形式。

For the language model example, since it just saw a subject, it might want to output information relevant to a verb, in case that’s what is coming next. For example, it might output whether the subject is singular or plural, so that we know what form a verb should be conjugated into if that’s what follows next.

长短时记忆的变体 Variants on Long Short Term Memory

到目前为止,我描述的是一个相当标准的 LSTM。但并非所有 LSTM 都与上述相同。事实上,几乎每篇涉及 LSTM 的论文都使用略有不同的版本。这些差异很小,但值得一提。

What I’ve described so far is a pretty normal LSTM. But not all LSTMs are the same as the above. In fact, it seems like almost every paper involving LSTMs uses a slightly different version. The differences are minor, but it’s worth mentioning some of them.

一个流行的 LSTM 变体由 Gers 和 Schmidhuber(2000)提出,即添加“窥视孔连接”。这意味着我们让门控层查看细胞状态。

One popular LSTM variant, introduced by Gers & Schmidhuber (2000), is adding “peephole connections.” This means that we let the gate layers look at the cell state.

上图对所有门都添加了窥视孔,但许多论文会只给部分门添加窥视孔。

The above diagram adds peepholes to all the gates, but many papers will give some peepholes and not others.

另一种变体是使用耦合的遗忘门和输入门。我们不再分别决定遗忘什么和添加什么新信息,而是将这些决策结合在一起。只有当我们打算在某个位置输入新内容时,我们才会遗忘;只有当我们遗忘一些旧内容时,我们才会向状态输入新值。

Another variation is to use coupled forget and input gates. Instead of separately deciding what to forget and what we should add new information to, we make those decisions together. We only forget when we’re going to input something in its place. We only input new values to the state when we forget something older.

一个更为显著的 LSTM 变体是门控循环单元(GRU),由 Cho 等人(2014)提出。它将遗忘门和输入门合并为一个单一的“更新门”。它还合并了细胞状态和隐藏状态,并进行了一些其他更改。得到的模型比标准 LSTM 模型更简单,并且越来越受欢迎。

A slightly more dramatic variation on the LSTM is the Gated Recurrent Unit, or GRU, introduced by Cho, et al. (2014). It combines the forget and input gates into a single “update gate.” It also merges the cell state and hidden state, and makes some other changes. The resulting model is simpler than standard LSTM models, and has been growing increasingly popular.

这些只是最著名的 LSTM 变体中的一部分。还有很多其他变体,例如 Yao 等人(2015)提出的深度门控 RNN。还有一些完全不同的方法来解决长期依赖问题,例如 Koutnik 等人(2014)提出的时钟频率 RNN。

These are only a few of the most notable LSTM variants. There are lots of others, like Depth Gated RNNs by Yao, et al. (2015). There’s also some completely different approach to tackling long-term dependencies, like Clockwork RNNs by Koutnik, et al. (2014).

这些变体中哪个最好?差异重要吗?Greff 等人(2015)对流行变体进行了很好的比较,发现它们几乎相同。Jozefowicz 等人(2015)测试了一万多种 RNN 架构,发现有些架构在某些任务上比 LSTM 表现更好。

Which of these variants is best? Do the differences matter? Greff, et al. (2015) do a nice comparison of popular variants, finding that they’re all about the same. Jozefowicz, et al. (2015) tested more than ten thousand RNN architectures, finding some that worked better than LSTMs on certain tasks.

结论 Conclusion

之前,我提到了人们使用 RNN 取得的显著成果。基本上所有这些成果都是使用 LSTM 实现的。对于大多数任务,它们确实要好得多!

Earlier, I mentioned the remarkable results people are achieving with RNNs. Essentially all of these are achieved using LSTMs. They really work a lot better for most tasks!

写成一组方程时,LSTM 看起来相当吓人。希望本文中逐步讲解的方式能让它们变得更容易理解一些。

Written down as a set of equations, LSTMs look pretty intimidating. Hopefully, walking through them step by step in this essay has made them a bit more approachable.

LSTM 是我们用 RNN 所能实现的一大进步。自然会想:还有下一个大进步吗?研究人员普遍认为:“有!下一步就是注意力机制!”其想法是让 RNN 的每一步从更大的信息集合中挑选信息来查看。例如,如果你用 RNN 为图像生成描述,它可能会为输出的每个单词选取图像的一部分来查看。事实上,Xu 等人(2015)正是这样做的——如果你想探索注意力机制,这可能是一个有趣的起点!使用注意力机制已经取得了一些非常令人兴奋的成果,而且似乎还有更多成果即将出现……

LSTMs were a big step in what we can accomplish with RNNs. It’s natural to wonder: is there another big step? A common opinion among researchers is: “Yes! There is a next step and it’s attention!” The idea is to let every step of an RNN pick information to look at from some larger collection of information. For example, if you are using an RNN to create a caption describing an image, it might pick a part of the image to look at for every word it outputs. In fact, Xu, _et al._ (2015) do exactly this – it might be a fun starting point if you want to explore attention! There’s been a number of really exciting results using attention, and it seems like a lot more are around the corner…

注意力机制并不是 RNN 研究中唯一令人兴奋的方向。例如,Kalchbrenner 等人(2015)的 Grid LSTM 看起来非常有前景。在生成模型中使用 RNN 的工作——如 Gregor 等人(2015)、Chung 等人(2015)或 Bayer & Osendorfer(2015)——也显得非常有趣。过去几年是循环神经网络的激动人心的时期,而未来几年只会更加精彩!

Attention isn’t the only exciting thread in RNN research. For example, Grid LSTMs by Kalchbrenner, _et al._ (2015) seem extremely promising. Work using RNNs in generative models – such as Gregor, _et al._ (2015), Chung, _et al._ (2015), or Bayer & Osendorfer (2015) – also seems very interesting. The last few years have been an exciting time for recurrent neural networks, and the coming ones promise to only be more so!

互动版:图/公式 + 针对本篇提问 →