A Recipe for Training Neural Networks
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→Andrej Karpathy 的博客:https://karpathy.github.io/2019/04/25/recipe/ 几周前,我发了一条关于“最常见的神经网络错误”的推文,列举了一些与训练神经网络相关的常见陷阱。这条推文获得的关注远超我的预期(包括一次网络研讨会)。显然,很多人都亲身经历过“卷积层是这样工作的”和“我们的卷积网络达到了最先进的水平”之间的巨大差距。因此,我想掸去博客上的灰尘,把我的推文扩展成这个主题应得的长文。然而,与其继续列举更多常见错误或详细阐述它们,我更想深入探讨如何从根本上避免这些错误(或快速修复它们)。诀窍在于遵循某个特定流程,据我所知,这个流程并没有被经常记录下来。让我们从两个重要的观察开始,正是它们引出了这个秘诀。
Andrej Karpathy's blog: https://karpathy.github.io/2019/04/25/recipe/ A few weeks ago, I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”. So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this topic deserves. However, instead of going into an enumeration of more common errors or fleshing them out, I wanted to dig a bit deeper and talk about how one can avoid making these errors altogether (or fix them very fast). The trick to doing so is to follow a certain process, which as far as I can tell is not very often documented. Let’s start with two important observations that motivate it.
Andrej Karpathy 的博客:https://karpathy.github.io/2019/04/25/recipe/
Andrej Karpathy's blog: https://karpathy.github.io/2019/04/25/recipe/
几周前,我在推特上发布了一条关于“最常见的神经网络错误”的推文,列举了训练神经网络时一些常见的陷阱。这条推文获得的关注比我预期的要多得多(还包括一次网络研讨会)。显然,很多人在“卷积层是如何工作的”和“我们的卷积网络达到了最先进的结果”之间亲身经历过巨大的鸿沟。
A few weeks ago, I posted a tweet on “the most common neural net mistakes”, listing a few common gotchas related to training neural nets. The tweet got quite a bit more engagement than I anticipated (including a webinar). Clearly, a lot of people have personally encountered the large gap between “here is how a convolutional layer works” and “our convnet achieves state of the art results”.
所以我想,是时候重拾我的博客,把这个主题以长文形式展开,因为它值得这样的篇幅。然而,与其列举更多的常见错误或详细阐述它们,我更想深入探讨如何完全避免这些错误(或非常快速地修复它们)。诀窍在于遵循一个特定的流程,据我所知,这个流程并不常见于文档。让我们从两个重要的观察开始,正是它们促使了该流程的形成。
So I thought it could be fun to brush off my dusty blog to expand my tweet to the long form that this topic deserves. However, instead of going into an enumeration of more common errors or fleshing them out, I wanted to dig a bit deeper and talk about how one can avoid making these errors altogether (or fix them very fast). The trick to doing so is to follow a certain process, which as far as I can tell is not very often documented. Let’s start with two important observations that motivate it.
据说,训练神经网络入门很容易。许多库和框架自豪地展示 30 行的神奇代码片段,可以解决你的数据问题,给人一种(错误的)印象,即这些东西是即插即用的。我们常常会看到类似这样的东西:
It is allegedly easy to get started with training neural nets. Numerous libraries and frameworks take pride in displaying 30-line miracle snippets that solve your data problems, giving the (false) impression that this stuff is plug and play. It’s common to see things like:
这些库和示例激活了我们大脑中熟悉标准软件的那部分——在那里,清晰的 API 和抽象往往是唾手可得的。以 Requests 库为例:
These libraries and examples activate the part of our brain that is familiar with standard software—a place where clean APIs and abstractions are often attainable. Consider the Requests library to demonstrate:
这很酷!一位勇敢的开发者已经替你承担了理解查询字符串、URL、GET/POST 请求、HTTP 连接等问题的负担,并将复杂性大体隐藏在几行代码之后。这是我们熟悉并期待的东西。遗憾的是,神经网络完全不是这样。只要你稍微偏离训练 ImageNet 分类器的范畴,它们就不再是“开箱即用”的技术。我曾在文章《是的,你应该理解反向传播》中试图指出这一点,拿反向传播开刀,称其为“有漏洞的抽象”,但现实情况其实还要糟糕得多。反向传播 + SGD 并不会神奇地让你的网络奏效;批归一化也不会神奇地让它收敛得更快;RNN 并不会神奇地让你“插入”文本;而且,仅仅因为你能把问题表述为强化学习(RL),并不意味着你就应该这么做。如果你坚持在不理解其原理的情况下使用这些技术,你很可能会失败。这就引出了……
That’s cool! A courageous developer has taken the burden of understanding query strings, URLs, GET/POST requests, HTTP connections, and so on from you and largely hidden the complexity behind a few lines of code. This is what we are familiar with and expect. Unfortunately, neural nets are nothing like that. They are not “off-the-shelf” technology the second you deviate slightly from training an ImageNet classifier. I’ve tried to make this point in my post “Yes you should understand backprop” by picking on backpropagation and calling it a “leaky abstraction”, but the situation is unfortunately much more dire. Backprop + SGD does not magically make your network work. Batch norm does not magically make it converge faster. RNNs don’t magically let you “plug in” text. And just because you can formulate your problem as RL doesn’t mean you should. If you insist on using the technology without understanding how it works you are likely to fail. Which brings me to…
当你破坏或错误配置代码时,通常会得到某种异常。例如,你在需要字符串的地方输入了整数;函数只期望 3 个参数;导入失败;键不存在;两个列表中的元素数量不相等。此外,通常可以为特定功能编写单元测试。
When you break or misconfigure code you will often get some kind of an exception. You plugged in an integer where something expected a string. The function only expected 3 arguments. This import failed. That key does not exist. The number of elements in the two lists is not equal. In addition, it is often possible to create unit tests for a certain functionality.
这仅仅是训练神经网络时的开始。一切在语法上可能是正确的,但整体安排不当,而且很难察觉。“可能的错误面”很大,是逻辑性(而非语法性)的,并且非常难以进行单元测试。例如,也许你在数据增强时对图像进行左右翻转,却忘记翻转标签。你的网络仍然可以(令人惊讶地)工作得相当好,因为网络可以在内部学习检测翻转图像,然后左右翻转其预测。或者,你的自回归模型由于差一错误(off-by-one bug)意外地将它试图预测的内容作为输入。或者你试图裁剪梯度,却裁剪了损失,导致训练时忽略离群样本。或者你从预训练检查点初始化权重,但没有使用原始均值。或者你只是弄错了正则化强度、学习率、其衰减率、模型大小等设置。因此,配置错误的神经网络只有在幸运时才会抛出异常;大多数时候它会训练,但静默地表现稍差。
This is just a start when it comes to training neural nets. Everything could be correct syntactically, but the whole thing is not arranged properly, and it is really hard to tell. The “possible error surface” is large, logical (as opposed to syntactic), and very tricky to unit test. For example, perhaps you forgot to flip your labels when you left-right flipped the image during data augmentation. Your net can still (shockingly) work pretty well because your network can internally learn to detect flipped images and then it left-right flips its predictions. Or maybe your autoregressive model accidentally takes the thing it is trying to predict as an input due to an off-by-one bug. Or you tried to clip your gradients but instead clipped the loss, causing the outlier examples to be ignored during training. Or you initialized your weights from a pre-trained checkpoint but did not use the original mean. Or you just screwed up the settings for regularization strengths, learning rate, its decay rate, model size, etc. Therefore, your misconfigured neural net will throw exceptions only if you are lucky; most of the time it will train but silently work a bit worse.
因此,(这一点再怎么强调也不为过)训练神经网络时,“快而猛”的方法是行不通的,只会带来痛苦。诚然,痛苦是让神经网络工作良好的一个非常自然的组成部分,但可以通过彻底、防御性、多疑以及痴迷于可视化几乎所有可能事物的态度来减轻。以我的经验,与深度学习成功最密切相关的品质是耐心和对细节的关注。
As a result (and this is really difficult to over-emphasize), a “fast and furious” approach to training neural networks does not work and only leads to suffering. Now, suffering is a perfectly natural part of getting a neural network to work well, but it can be mitigated by being thorough, defensive, paranoid, and obsessed with visualizations of basically every possible thing. The qualities that in my experience correlate most strongly to success in deep learning are patience and attention to detail.
鉴于上述两个事实,我为自己制定了一套具体流程,在将神经网络应用于新问题时会遵循,下面我将尝试描述它。你会看到它非常重视上述两条原则。特别是,它从简单到复杂逐步构建,并且在每一步我们都会对将要发生的事情做出具体假设,然后用实验加以验证,或者调查下去直到发现问题。我们极力避免的是同时引入大量“未经验证”的复杂性,这必然会带来 bug/配置错误,而这些错误将花费很长时间才能找到(如果能找到的话)。如果编写神经网络代码就像训练神经网络一样,你会希望使用非常小的学习率,并在每次迭代后猜测并评估完整测试集。
In light of the above two facts, I have developed a specific process for myself that I follow when applying a neural net to a new problem, which I will try to describe. You will see that it takes the two principles above very seriously. In particular, it builds from simple to complex and at every step of the way we make concrete hypotheses about what will happen and then either validate them with an experiment or investigate until we find some issue. What we try to prevent very hard is the introduction of a lot of “unverified” complexity at once, which is bound to introduce bugs/misconfigurations that will take forever to find (if ever). If writing your neural net code was like training one, you’d want to use a very small learning rate and guess and then evaluate the full test set after every iteration.
训练神经网络的第一步是完全不碰任何神经网络代码,而是从彻底检查数据开始。这一步至关重要。我喜欢花大量时间(以小时计)浏览成千上万个样本,理解它们的分布并寻找模式。幸运的是,你的大脑很擅长这件事。有一次我发现数据中包含重复样本;另一次我找到了损坏的图像/标签。我会关注数据不平衡和偏差。我通常还会留意自己给数据分类的过程,这能暗示我们最终可能要探索的架构类型。例如——局部特征是否足够,还是需要全局上下文?数据有多少变化,以什么形式出现?哪些变化是虚假的、可以通过预处理去除?空间位置是否重要,还是我们希望做平均池化将其去除?细节有多重要,我们能承受将图像下采样到什么程度?标签噪声有多大?
The first step to training a neural net is to not touch any neural net code at all and instead begin by thoroughly inspecting your data. This step is critical. I like to spend copious amounts of time (measured in units of hours) scanning through thousands of examples, understanding their distribution and looking for patterns. Luckily, your brain is pretty good at this. One time I discovered that the data contained duplicate examples. Another time I found corrupted images / labels. I look for data imbalances and biases. I will typically also pay attention to my own process for classifying the data, which hints at the kinds of architectures we’ll eventually explore. As an example - are very local features enough or do we need global context? How much variation is there and what form does it take? What variation is spurious and could be preprocessed out? Does spatial position matter or do we want to average pool it out? How much does detail matter and how far could we afford to downsample the images? How noisy are the labels?
此外,由于神经网络实际上是你数据集的压缩/编译版本,你可以观察网络的(错误)预测,并理解它们可能来自何处。如果网络给出的某个预测与你在数据中看到的不一致,那说明有问题。
In addition, since the neural net is effectively a compressed/compiled version of your dataset, you’ll be able to look at your network (mis)predictions and understand where they might be coming from. And if your network is giving you some prediction that doesn’t seem consistent with what you’ve seen in the data, something is off.
在获得定性认识之后,最好再写一些简单的代码,按你能想到的任何维度(如标签类型、标注大小、标注数量等)进行搜索/过滤/排序,并将它们的分布以及任何轴上的离群点可视化。离群点尤其几乎总能暴露出数据质量或预处理中的某些 bug。
Once you get a qualitative sense it is also a good idea to write some simple code to search/filter/sort by whatever you can think of (e.g. type of label, size of annotations, number of annotations, etc.) and visualize their distributions and the outliers along any axis. The outliers especially almost always uncover some bugs in data quality or preprocessing.
既然我们理解了数据,我们能否直接拿起超炫酷的多尺度 ASPP FPN ResNet 并开始训练出色的模型?当然不能。那是通往痛苦之路。我们的下一步是搭建一个完整的训练+评估骨架,并通过一系列实验来信任其正确性。在这个阶段,最好选择一个你不可能搞砸的简单模型——例如线性分类器或非常小的 ConvNet。我们需要训练它,可视化损失、任何其他指标(如准确率)、模型预测,并一路进行一系列带有明确假设的消融实验。
Now that we understand our data can we reach for our super fancy Multi-scale ASPP FPN ResNet and begin training awesome models? For sure no. That is the road to suffering. Our next step is to set up a full training + evaluation skeleton and gain trust in its correctness via a series of experiments. At this stage it is best to pick some simple model that you couldn’t possibly have screwed up somehow - e.g. a linear classifier, or a very tiny ConvNet. We’ll want to train it, visualize the losses, any other metrics (e.g. accuracy), model predictions, and perform a series of ablation experiments with explicit hypotheses along the way.
固定随机种子。始终使用固定的随机种子,以保证你两次运行代码时得到相同的结果。这消除了一个变化因素,并帮助你保持清醒。
fix random seed. Always use a fixed random seed to guarantee that when you run the code twice you will get the same outcome. This removes a factor of variation and will help keep you sane.
简化。确保禁用任何不必要的花哨功能。例如,在这个阶段一定要关闭任何数据增强。数据增强是一种正则化策略,我们之后可能会采用,但现在它只是引入愚蠢错误的又一个机会。
simplify. Make sure to disable any unnecessary fanciness. As an example, definitely turn off any data augmentation at this stage. Data augmentation is a regularization strategy that we may incorporate later, but for now it is just another opportunity to introduce some dumb bug.
在评估中加入有效数字。在绘制测试损失时,要在整个(大)测试集上运行评估。不要仅仅按批次绘制测试损失,然后依赖 Tensorboard 中的平滑。我们追求的是正确性,并且非常愿意为了保持清醒而放弃时间。
add significant digits to your eval. When plotting the test loss run the evaluation over the entire (large) test set. Do not just plot test losses over batches and then rely on smoothing them in Tensorboard. We are in pursuit of correctness and are very willing to give up time for staying sane.
验证初始化时的损失。验证你的损失从正确的损失值开始。例如,如果正确初始化最后一层,你应该测量到
verify loss @ init. Verify that your loss starts at the correct loss value. E.g. if you initialize your final layer correctly you should measure
在初始化时对 softmax 层进行设置。同样的默认值也可以推导出来,用于 L2 回归、Huber 损失等。
on a softmax at initialization. The same default values can be derived for L2 regression, Huber losses, etc.
* 初始化要好。正确初始化最后一层的权重。例如,如果你回归一些均值为 50 的数值,那么将最后的偏置初始化为 50。如果你的数据集正负样本比是 1:10 的不平衡数据,则设置 logits 的偏置,使网络在初始化时预测概率为 0.1。正确设置这些参数会加快收敛,并消除“曲棍球棒”形状的损失曲线——在最初几次迭代中,网络基本上只是在学习偏置。
* init well. Initialize the final layer weights correctly. E.g. if you are regressing some values that have a mean of 50 then initialize the final bias to 50. If you have an imbalanced dataset of a ratio 1:10 of positives:negatives, set the bias on your logits such that your network predicts probability of 0.1 at initialization. Setting these correctly will speed up convergence and eliminate “hockey stick” loss curves where in the first few iteration your network is basically just learning the bias.
* 人类基线。监控除损失之外人类可解释且可检查的指标(如准确率)。只要可能,就评估你自己(人类)的准确率并与之比较。或者,对测试数据标注两次,将每次标注分别作为预测和真实标签。
* human baseline. Monitor metrics other than loss that are human interpretable and checkable (e.g. accuracy). Whenever possible evaluate your own (human) accuracy and compare to it. Alternatively, annotate the test data twice and for each example treat one annotation as prediction and the second as ground truth.
* 输入无关基线。训练一个输入无关的基线(例如最简单的做法是把所有输入都设为零)。它的表现应该比你真正使用数据(不置零)时更差。是这样吗?也就是说,你的模型是否真的从输入中提取到了信息?
* input-independent baseline. Train an input-independent baseline, (e.g. easiest is to just set all your inputs to zero). This should perform worse than when you actually plug in your data without zeroing it out. Does it? i.e. does your model learn to extract any information out of the input at all?
* 过拟合一个批次。在一个只有少量样本的批次上过拟合(例如最少两个样本)。为此,我们增大模型的容量(例如增加层或滤波器),并验证我们能够达到可实现的最小损失(例如零)。我还喜欢在同一张图中同时画出标签和预测,并确保在达到最小损失时它们完全吻合。如果没有,说明某处存在 bug,我们就不能继续下一阶段。
* overfit one batch. Overfit a single batch of only a few examples (e.g. as little as two). To do so we increase the capacity of our model (e.g. add layers or filters) and verify that we can reach the lowest achievable loss (e.g. zero). I also like to visualize in the same plot both the label and the prediction and ensure that they end up aligning perfectly once we reach the minimum loss. If they do not, there is a bug somewhere and we cannot continue to the next stage.
* 验证训练损失是否下降。在这个阶段,你很可能因为使用的是一个玩具模型而在数据集上欠拟合。尝试稍微增加其容量。你的训练损失是否如预期那样下降了?
* verify decreasing training loss. At this stage you will hopefully be underfitting on your dataset because you’re working with a toy model. Try to increase its capacity just a bit. Did your training loss go down as it should?
* 就在网络之前进行可视化。可视化数据的唯一正确位置,就是紧挨在你的网络之前(在 tf 中)。
* visualize just before the net. The unambiguously correct place to visualize your data is immediately before your net (in tf).
也就是说——你要把进入网络的**确切**内容可视化,把原始数据张量和标签解码成可视化结果。这是唯一的“真相来源”。我数不清有多少次这救了我,并揭示出数据预处理和增强中的问题。
That is - you want to visualize _exactly_ what goes into your network, decoding that raw tensor of data and labels into visualizations. This is the only “source of truth”. I can’t count the number of times this has saved me and revealed problems in data preprocessing and augmentation.
* 可视化预测动态。我喜欢在训练过程中对固定测试批次可视化模型的预测。这些预测移动的“动态”将让你对训练进展产生非常好的直觉。很多时候,如果网络以某种方式过度摆动,你就能感觉到它在“挣扎”着拟合数据,从而暴露出不稳定性。学习率过低或过高也很容易从抖动幅度中察觉。
* visualize prediction dynamics. I like to visualize model predictions on a fixed test batch during the course of training. The “dynamics” of how these predictions move will give you incredibly good intuition for how the training progresses. Many times it is possible to feel the network “struggle” to fit your data if it wiggles too much in some way, revealing instabilities. Very low or very high learning rates are also easily noticeable in the amount of jitter.
* 使用反向传播来绘制依赖关系。你的深度学习代码通常会包含复杂的、向量化的和广播式的操作。我遇到过几次的一个相对常见的错误是,人们在这里出错(例如,他们使用
* use backprop to chart dependencies. Your deep learning code will often contain complicated, vectorized, and broadcasted operations. A relatively common bug I’ve come across a few times is that people get this wrong (e.g. they use
某处) 并意外地在批次维度上混合了信息。令人沮丧的是,你的网络通常仍然可以正常训练,因为它会学会忽略来自其他示例的数据。调试此问题(以及其他相关问题)的一种方法是将损失设置为一些琐碎的东西,例如示例 i 的所有输出之和,将反向传播一直运行到输入,并确保只在第 i 个输入上获得非零梯度。同样的策略也可用于例如确保你的自回归模型在时间 t 只依赖于 1..t-1。更一般地,梯度为你提供关于网络中什么依赖于什么的信息,这对调试很有用。
somewhere) and inadvertently mix information across the batch dimension. It is a depressing fact that your network will typically still train okay because it will learn to ignore data from the other examples. One way to debug this (and other related problems) is to set the loss to be something trivial like the sum of all outputs of example i, run the backward pass all the way to the input, and ensure that you get a non-zero gradient only on the i-th input. The same strategy can be used to e.g. ensure that your autoregressive model at time t only depends on 1..t-1. More generally, gradients give you information about what depends on what in your network, which can be useful for debugging.
* 泛化一个特例。这更像是一个通用的编码技巧,但我经常看到人们在试图一口吃成胖子时制造 bug,他们从头编写一个相对通用的功能。我喜欢为当前正在做的事情编写一个非常具体的函数,让它工作,然后再泛化它,确保得到相同的结果。这通常适用于向量化代码,我几乎总是先写出完整的循环版本,然后才一次一个循环地将其转换为向量化代码。
* generalize a special case. This is a bit more of a general coding tip but I’ve often seen people create bugs when they bite off more than they can chew, writing a relatively general functionality from scratch. I like to write a very specific function to what I’m doing right now, get that to work, and then generalize it later making sure that I get the same result. Often this applies to vectorizing code, where I almost always write out the fully loopy version first and only then transform it to vectorized code one loop at a time.
在这个阶段,我们应该对数据集有很好的理解,并且已经有完整的训练和评估流程在工作。对于任何给定的模型,我们都可以(可复现地)计算一个我们信任的指标。我们还具备输入无关基线的性能、几个简单基线的性能(我们最好超过这些),并且我们粗略了解人类的性能(我们希望达到这个水平)。现在,我们已经准备好迭代出好的模型。
At this stage we should have a good understanding of the dataset and we have the full training + evaluation pipeline working. For any given model we can (reproducibly) compute a metric that we trust. We are also armed with our performance for an input-independent baseline, the performance of a few dumb baselines (we better beat these), and we have a rough sense of the performance of a human (we hope to reach this). The stage is now set for iterating on a good model.
我喜欢用两阶段方法来寻找好的模型:首先获得一个足够大的模型,使其能够过拟合(即关注训练损失),然后适当地正则化(放弃一些训练损失以改善验证损失)。我之所以喜欢这两个阶段,是因为如果我们根本无法用任何模型达到较低的误差率,那可能再次表明存在某些问题、错误或配置不当。
The approach I like to take to finding a good model has two stages: first get a model large enough that it can overfit (i.e. focus on training loss) and then regularize it appropriately (give up some training loss to improve the validation loss). The reason I like these two stages is that if we are not able to reach a low error rate with any model at all that may again indicate some issues, bugs, or misconfiguration.
* 选择模型。要获得良好的训练损失,你需要为数据选择合适的架构。关于如何选择,我的首要建议是:别逞英雄。我见过很多人渴望发挥创造力,用各种奇特的架构搭建神经网络工具箱的积木,这些架构在他们看来很有道理。在项目早期阶段要强烈抵制这种诱惑。我总是建议人们直接找到最相关的论文,复制粘贴其中实现良好性能的最简单架构。例如,如果你在分类图像,别逞英雄,第一次运行直接复制粘贴 ResNet-50。以后你可以做更定制化的东西并超越它。
* picking the model. To reach a good training loss you’ll want to choose an appropriate architecture for the data. When it comes to choosing this my #1 advice is: Don’t be a hero. I’ve seen a lot of people who are eager to get crazy and creative in stacking up the lego blocks of the neural net toolbox in various exotic architectures that make sense to them. Resist this temptation strongly in the early stages of your project. I always advise people to simply find the most related paper and copy paste their simplest architecture that achieves good performance. E.g. if you are classifying images don’t be a hero and just copy paste a ResNet-50 for your first run. You’re allowed to do something more custom later and beat this.
* Adam 是安全的。在设定基线的早期阶段,我喜欢使用学习率为 3e-4 的 Adam。根据我的经验,Adam 对超参数(包括糟糕的学习率)更加宽容。对于卷积神经网络(ConvNets),经过良好调优的 SGD 几乎总是比 Adam 略胜一筹,但最优学习率范围要窄得多,而且因问题而异。(注意:如果你使用 RNN 和相关序列模型,使用 Adam 更为常见。在项目初期,再次强调,别逞英雄,跟随最相关论文的做法。)
* adam is safe. In the early stages of setting baselines I like to use Adam with a learning rate of 3e-4. In my experience Adam is much more forgiving to hyperparameters, including a bad learning rate. For ConvNets a well-tuned SGD will almost always slightly outperform Adam, but the optimal learning rate region is much more narrow and problem-specific. (Note: If you are using RNNs and related sequence models it is more common to use Adam. At the initial stage of your project, again, don’t be a hero and follow whatever the most related papers do.)
* 一次只增加一个复杂度。如果你有多个信号要接入分类器,我建议你一个一个地接入,并且每次确保获得你期望的性能提升。不要在开始时把所有东西都一股脑塞给模型。还有其他构建复杂性的方法——例如,你可以先尝试输入较小的图像,然后再增大尺寸,等等。
* complexify only one at a time. If you have multiple signals to plug into your classifier I would advise that you plug them in one by one and every time ensure that you get a performance boost you’d expect. Don’t throw the kitchen sink at your model at the start. There are other ways of building up complexity - e.g. you can try to plug in smaller images first and make them bigger later, etc.
不要轻信学习率衰减的默认设置。如果你从其他领域复用代码,务必对学习率衰减格外小心。不仅不同问题需要不同的衰减调度,而且更糟糕的是,在典型实现中,调度将基于当前的 epoch 数,而 epoch 数仅仅因数据集大小不同就可能差异很大。例如,ImageNet 会在第 30 个 epoch 时将学习率衰减 10 倍。如果你不是在训练 ImageNet,那么几乎可以肯定你不需要这个设置。如果不小心,你的代码可能会过早地把学习率秘密地推向零,从而阻止模型收敛。就我自己的工作而言,我总是完全禁用学习率衰减(使用恒定学习率),并在最后阶段才对其进行调节。
Do not trust learning rate decay defaults. If you are re-purposing code from some other domain, always be very careful with learning rate decay. Not only would you want to use different decay schedules for different problems, but—even worse—in a typical implementation the schedule will be based on the current epoch number, which can vary widely simply depending on the size of your dataset. For example, ImageNet would decay by 10 on epoch 30. If you're not training ImageNet, then you almost certainly do not want this. If you're not careful, your code could secretly be driving your learning rate to zero too early, not allowing your model to converge. In my own work, I always disable learning rate decays entirely (I use a constant LR) and tune this all the way at the very end.
理想情况下,我们现在处于这样一个阶段:我们有一个大模型,至少能够拟合训练集。现在是时候对其进行正则化,通过牺牲部分训练准确率来获得一些验证准确率。以下是一些技巧与建议:
Ideally, we are now at a place where we have a large model that is fitting at least the training set. Now it is time to regularize it and gain some validation accuracy by giving up some of the training accuracy. Some tips & tricks:
* 获取更多数据。首先,在任何实际场景中,对模型进行正则化的最佳且首选方式始终是添加更多真实训练数据。一个非常常见的错误是花费大量工程精力试图从一个小数据集中榨取收益,而其实你可以转而收集更多数据。据我所知,添加更多数据几乎是唯一能保证让一个配置良好的神经网络几乎无限期地单调提升性能的方法。另一种方法是集成(如果你能负担得起的话),但集成在大约 5 个模型之后就达到上限了。
* Get more data. First, the by far best and preferred way to regularize a model in any practical setting is to add more real training data. It is a very common mistake to spend a lot of engineering cycles trying to squeeze juice out of a small dataset when you could instead be collecting more data. As far as I'm aware, adding more data is pretty much the only guaranteed way to monotonically improve the performance of a well-configured neural network almost indefinitely. The other would be ensembles (if you can afford them), but that tops out after ~5 models.
* 数据增强。仅次于真实数据的最佳选择是半合成数据——尝试更激进的数据增强方法。
* Data augmentation. The next best thing to real data is half-fake data - try out more aggressive data augmentation.
* 创造性增强。如果半合成数据不起作用,合成数据也可能有所作为。人们正在寻找创造性的方法来扩充数据集;例如,域随机化、使用仿真、巧妙的混合方法(如将(可能经仿真的)数据插入场景中),甚至使用 GANs。
* Creative augmentation. If half-fake data doesn't do it, fake data may also do something. People are finding creative ways of expanding datasets; for example, domain randomization, use of simulation, clever hybrids such as inserting (potentially simulated) data into scenes, or even GANs.
* 预训练。如果可能的话,使用预训练网络几乎从不会有坏处,即使你有足够的数据。
* Pretrain. It rarely ever hurts to use a pretrained network if you can, even if you have enough data.
- 坚持使用监督学习。不要对无监督预训练过于兴奋。与 2008 年那篇博客文章告诉你的不同,据我所知,在当代计算机视觉中,还没有任何版本的无监督预训练报告过强劲的结果(尽管 NLP 近年来在 BERT 及其同类模型上表现不错,这很可能归因于文本更具刻意性,以及更高的信噪比)。
- Stick with supervised learning. Do not get over-excited about unsupervised pretraining. Unlike what that blog post from 2008 tells you, as far as I know, no version of it has reported strong results in modern computer vision (though NLP seems to be doing pretty well with BERT and friends these days, quite likely owing to the more deliberate nature of text, and a higher signal-to-noise ratio).
- 降低输入维度。移除可能包含伪信号的特征。如果数据集很小,任何增加的伪输入都只是另一次过拟合的机会。类似地,如果低级细节不太重要,尝试输入较小的图像。
- Smaller input dimensionality. Remove features that may contain spurious signal. Any added spurious input is just another opportunity to overfit if your dataset is small. Similarly, if low-level details don't matter much, try to input a smaller image.
- 缩小模型规模。在许多情况下,你可以利用领域知识对网络施加约束来减小其规模。例如,过去流行在 ImageNet 的骨干网络顶部使用全连接层,但后来这些层被简单的平均池化所取代,从而消除了大量参数。
- Smaller model size. In many cases you can use domain knowledge constraints on the network to decrease its size. As an example, it used to be trendy to use fully connected layers at the top of backbones for ImageNet, but these have since been replaced with simple average pooling, eliminating a ton of parameters in the process.
- 减小批大小。由于批归一化内部的归一化操作,较小的批大小在一定程度上对应更强的正则化。这是因为批经验均值/标准差是整体均值/标准差的更近似版本,因此缩放和偏移会更多地“摆动”你的批次。
- Decrease the batch size. Due to the normalization inside batch norm, smaller batch sizes somewhat correspond to stronger regularization. This is because the batch empirical mean/std are more approximate versions of the full mean/std, so the scale & offset “wiggles” your batch around more.
- Dropout。添加 dropout。对于卷积网络使用 dropout2d(空间 dropout)。要谨慎/小心地使用它,因为 dropout 似乎与批归一化不太兼容。
- Drop. Add dropout. Use dropout2d (spatial dropout) for ConvNets. Use this sparingly/carefully because dropout does not seem to play nice with batch normalization.
* **权重衰减。** 增加权重衰减惩罚。
* **Weight decay.** Increase the weight decay penalty.
* **早停。** 根据测得的验证损失停止训练,以便在模型即将过拟合时将其捕捉。
* **Early stopping.** Stop training based on your measured validation loss to catch your model just as it’s about to overfit.
* **尝试更大的模型。** 我最后才提到这一点,并且只有在早停之后才提到,但过去我多次发现,更大的模型最终当然会过拟合得多,但它们的“早停”性能往往比小模型好得多。
* **Try a larger model.** I mention this last and only after early stopping but I’ve found a few times in the past that larger models will of course overfit much more eventually, but their “early stopped” performance can often be much better than that of smaller models.
最后,为了进一步确信你的网络是一个合理的分类器,我喜欢可视化网络的第一层权重,并确保得到有意义的漂亮边缘。如果你的第一层滤波器看起来像噪声,那么可能有问题。同样,网络内部的激活有时会显示奇怪的伪影并暗示问题。
Finally, to gain additional confidence that your network is a reasonable classifier, I like to visualize the network’s first-layer weights and ensure you get nice edges that make sense. If your first layer filters look like noise then something could be off. Similarly, activations inside the net can sometimes display odd artifacts and hint at problems.
你现在应该已经“进入循环”,与你的数据集一起,广泛探索能够实现低验证损失的模型架构。本步骤的一些技巧与提示:
You should now be “in the loop” with your dataset exploring a wide model space for architectures that achieve low validation loss. A few tips and tricks for this step:
* **用随机搜索代替网格搜索**。在同时调整多个超参数时,使用网格搜索以确保覆盖所有设置听起来很有吸引力,但请记住,最好改用随机搜索。直觉上,这是因为神经网络通常对某些参数比对其他参数更敏感。在极端情况下,如果参数 *a* 很重要而改变 *b* 没有影响,那么你宁愿更彻底地对 *a* 进行采样,而不是在几个固定点上重复多次。
* Random search over grid search. For simultaneously tuning multiple hyperparameters, it may sound tempting to use grid search to ensure coverage of all settings, but keep in mind that it is best to use random search instead. Intuitively, this is because neural nets are often much more sensitive to some parameters than others. In the limit, if a parameter *a* matters but changing *b* has no effect, then you would rather sample *a* more thoroughly than at a few fixed points multiple times.
* **超参数优化**。市面上有大量花哨的贝叶斯超参数优化工具箱,我的几位朋友也报告说它们效果不错,但就我个人经验而言,探索一个漂亮且广泛的模型与超参数空间的最先进方法是使用一名实习生 :)。开个玩笑。
* Hyper-parameter optimization. There is a large number of fancy Bayesian hyper-parameter optimization toolboxes around and a few of my friends have also reported success with them, but my personal experience is that the state-of-the-art approach to exploring a nice and wide space of models and hyperparameters is to use an intern :). Just kidding.
一旦你找到了最佳的架构类型和超参数,你仍然可以使用一些额外的技巧来榨取系统的最后一点性能: * **集成模型。** 模型集成几乎是一种保证能在任何任务上提升 2%准确率的方法。如果你无法承受测试时的计算开销,可以考虑利用暗知识将集成模型蒸馏到一个网络中。
Once you find the best types of architectures and hyperparameters you can still use a few more tricks to squeeze out the last pieces of juice out of the system: * **Ensembles.** Model ensembles are a pretty much guaranteed way to gain 2% of accuracy on anything. If you can’t afford the computation at test time look into distilling your ensemble into a network using dark knowledge.
* **让它继续训练。** 我经常看到人们倾向于在验证损失似乎趋于平稳时停止模型训练。根据我的经验,网络会持续训练很长一段时间,这超出了直觉。有一次,我意外地让一个模型在寒假期间继续训练,等我一月份回来时,它已经达到了 SOTA(“最先进水平”)。
* **Leave it training.** I’ve often seen people tempted to stop the model training when the validation loss seems to be leveling off. In my experience networks keep training for unintuitively long time. One time I accidentally left a model training during the winter break and when I got back in January it was SOTA (“state of the art”).
Once you find the best types of architectures and hyper-parameters you can still use a few more tricks to squeeze out the last pieces of juice out of the system:
一旦你走到这一步,你就拥有了成功的所有要素:你对技术、数据集和问题有了深刻的理解,你建立了完整的训练/评估基础设施,并对其准确性充满信心,而且你探索了越来越复杂的模型,并按照你每一步预测的方式获得了性能提升。现在你已准备好阅读大量论文、尝试大量实验,并取得你的 SOTA(最先进)结果。祝你好运!
Once you make it here you’ll have all the ingredients for success: You have a deep understanding of the technology, the dataset and the problem, you’ve set up the entire training/evaluation infrastructure and achieved high confidence in its accuracy, and you’ve explored increasingly more complex models, gaining performance improvements in ways you’ve predicted each step of the way. You’re now ready to read a lot of papers, try a large number of experiments, and get your SOTA results. Good luck!