Software 2.0
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→我有时看到人们将神经网络仅仅视为“机器学习工具箱中的另一个工具”。它们有优点和缺点,在某些地方有效,有时你可以用它们赢得 Kaggle 竞赛。不幸的是,这种解释完全只见树木不见森林。神经网络不仅仅是另一个分类器,它们代表了我们开发软件方式的根本性转变。它们是软件 2.0。我们熟悉的“经典栈”软件 1.0 是用 Python、C++ 等语言编写的。它由程序员编写的对计算机的显式指令组成。通过编写每一行代码,程序员在程序空间中确定一个具有某些期望行为的特定点。
[](https://karpathy.medium.com/?source=post_page---byline--a64152b37c35---------------------------------------) I sometimes see people refer to neural networks as just “another tool in your machine learning toolbox”. They have some pros and cons, they work here or there, and sometimes you can use them to win Kaggle competitions. Unfortunately, this interpretation completely misses the forest for the trees. Neural networks are not just another classifier, they represent the beginning of a fundamental shift in how we develop software. They are Software 2.0. The “classical stack” of Software 1.0 is what we’re all familiar with — it is written in languages such as Python, C++, etc. It consists of explicit instructions to the computer written by a programmer. By writing each line of code, the programmer identifies a specific point in program space with some desirable behavior.
[](https://karpathy.medium.com/?source=post_page---byline--a64152b37c35---------------------------------------)
[](https://karpathy.medium.com/?source=post_page---byline--a64152b37c35---------------------------------------)
我有时看到人们将神经网络称为“机器学习工具箱中的又一个工具”。它们有优点也有缺点,在某些地方有效,有时你可以用它们赢得 Kaggle 竞赛。不幸的是,这种解释完全只见树木不见森林。神经网络不仅仅是另一个分类器,它们代表了我们开发软件方式的根本性转变的开端。它们是软件 2.0。
I sometimes see people refer to neural networks as just “another tool in your machine learning toolbox”. They have some pros and cons, they work here or there, and sometimes you can use them to win Kaggle competitions. Unfortunately, this interpretation completely misses the forest for the trees. Neural networks are not just another classifier, they represent the beginning of a fundamental shift in how we develop software. They are Software 2.0.
软件 1.0 的“经典栈”是我们都熟悉的——它用 Python、C++等语言编写。它由程序员编写的对计算机的显式指令组成。通过编写每一行代码,程序员在程序空间中确定一个具有某种期望行为的具体点。
The “classical stack” of Software 1.0 is what we’re all familiar with — it is written in languages such as Python, C++, etc. It consists of explicit instructions to the computer written by a programmer. By writing each line of code, the programmer identifies a specific point in program space with some desirable behavior.
按回车或点击查看全尺寸图片
Press enter or click to view image in full size
相比之下,软件 2.0 是用更抽象、对人类不友好的语言编写的,例如神经网络的权重。没有人类参与编写这些代码,因为权重很多(典型网络可能有数百万个),而且直接用权重编码有点难(我试过)。
In contrast, Software 2.0 is written in much more abstract, human unfriendly language, such as the weights of a neural network. No human is involved in writing this code because there are a lot of weights (typical networks might have millions), and coding directly in weights is kind of hard (I tried).
按回车或点击查看全尺寸图片
Press enter or click to view image in full size
相反,我们的方法是指定一个期望程序行为的某些目标(例如,“满足一个输入输出对示例的数据集”,或“赢得一局围棋”),编写代码的粗略骨架(即神经网络架构),该骨架标识出要搜索的程序空间子集,然后利用我们可用的计算资源在这个空间中搜索一个有效的程序。在神经网络的情况下,我们将搜索限制在程序空间的连续子集上,在这个子集中,搜索过程可以通过反向传播和随机梯度下降(有些令人惊讶地)变得高效。
Instead, our approach is to specify some goal on the behavior of a desirable program (e.g., “satisfy a dataset of input output pairs of examples”, or “win a game of Go”), write a rough skeleton of the code (i.e. a neural net architecture) that identifies a subset of program space to search, and use the computational resources at our disposal to search this space for a program that works. In the case of neural networks, we restrict the search to a continuous subset of the program space where the search process can be made (somewhat surprisingly) efficient with backpropagation and stochastic gradient descent.
按回车或点击查看全尺寸图片
Press enter or click to view image in full size
为了明确类比,在软件 1.0 中,人工编写的源代码(例如一些.cpp 文件)被编译成执行有用工作的二进制文件。在软件 2.0 中,源代码通常包括:1)定义期望行为的数据集,以及 2)给出代码粗略骨架的神经网络架构,但有许多细节(权重)需要填充。训练神经网络的过程将数据集编译成二进制文件——最终的神经网络。在当今的大多数实际应用中,神经网络架构和训练系统越来越标准化为一种商品,因此大部分活跃的“软件开发”以策划、增长、处理和清理标记数据集的形式进行。这从根本上改变了我们迭代软件的编程范式,因为团队分为两部分:2.0 程序员(数据标注员)编辑和增长数据集,而少数 1.0 程序员维护和迭代周围的训练代码基础设施、分析、可视化和标注界面。
To make the analogy explicit, in Software 1.0, human-engineered source code (e.g. some .cpp files) is compiled into a binary that does useful work. In Software 2.0 most often the source code comprises 1) the dataset that defines the desirable behavior and 2) the neural net architecture that gives the rough skeleton of the code, but with many details (the weights) to be filled in. The process of training the neural network compiles the dataset into the binary — the final neural network. In most practical applications today, the neural net architectures and the training systems are increasingly standardized into a commodity, so most of the active “software development” takes the form of curating, growing, massaging and cleaning labeled datasets. This is fundamentally altering the programming paradigm by which we iterate on our software, as the teams split in two: the 2.0 programmers (data labelers) edit and grow the datasets, while a few 1.0 programmers maintain and iterate on the surrounding training code infrastructure, analytics, visualizations and labeling interfaces.
事实证明,现实世界中的很大一部分问题具有这样的特性:收集数据(或更一般地,确定期望行为)比显式编写程序要容易得多。由于这一点以及下面我将讨论的软件 2.0 程序的许多其他好处,我们正在目睹整个行业的大规模转型,大量 1.0 代码正在被移植到 2.0 代码。软件(1.0)正在吞噬世界,而现在 AI(软件 2.0)正在吞噬软件。
It turns out that a large portion of real-world problems have the property that it is significantly easier to collect the data (or more generally, identify a desirable behavior) than to explicitly write the program. Because of this and many other benefits of Software 2.0 programs that I will go into below, we are witnessing a massive transition across the industry where of a lot of 1.0 code is being ported into 2.0 code. Software (1.0) is eating the world, and now AI (Software 2.0) is eating software.
让我们简要审视一些正在进行的转变的具体例子。在以下每个领域中,我们都看到了过去几年中的改进,当我们放弃尝试通过编写显式代码来解决复杂问题,而是将代码过渡到 2.0 堆栈时。
Let’s briefly examine some concrete examples of this ongoing transition. In each of these areas we’ve seen improvements over the last few years when we give up on trying to address a complex problem by writing explicit code and instead transition the code into the 2.0 stack.
视觉识别曾经由手工设计的特征组成,最后再点缀一点机器学习(例如 SVM)。自那以后,我们通过获取大型数据集(如 ImageNet)并在卷积神经网络架构空间中搜索,发现了更强大的视觉特征。最近,我们甚至不再信任自己手工设计架构,而是开始也对这些架构进行搜索。
Visual Recognition used to consist of engineered features with a bit of machine learning sprinkled on top at the end (e.g., an SVM). Since then, we discovered much more powerful visual features by obtaining large datasets (e.g. ImageNet) and searching in the space of Convolutional Neural Network architectures. More recently, we don’t even trust ourselves to hand-code the architectures and we’ve begun searching over those as well.
语音识别曾经涉及大量预处理、高斯混合模型和隐马尔可夫模型,但如今几乎完全由神经网络组成。一个非常相关且常被引用的幽默语录来自 1985 年的 Fred Jelinek:“每当我解雇一位语言学家,我们语音识别系统的性能就会提升。”
Speech recognition used to involve a lot of preprocessing, gaussian mixture models and hidden markov models, but todayconsist almost entirely of neural net stuff. A very related, often cited humorous quote attributed to Fred Jelinek from 1985 reads “Every time I fire a linguist, the performance of our speech recognition system goes up”.
语音合成历史上曾采用各种拼接机制,但如今最先进的模型是生成原始音频信号输出的大型卷积网络(如 WaveNet)。
Speech synthesis has historically been approached with various stitching mechanisms, but today the state of the art models are large ConvNets (e.g. WaveNet) that produce raw audio signal outputs.
机器翻译通常采用基于短语的统计技术,但神经网络正迅速占据主导地位。我最喜欢的架构是在多语言设置中训练的,其中单个模型从任何源语言翻译到任何目标语言,并且在弱监督(或完全无监督)设置中。
Machine Translation has usually been approaches with phrase-based statistical techniques, but neural networks are quickly becoming dominant. My favorite architectures are trained in the multilingual setting, where a single model translates from any source language to any target language, and in weakly supervised (or entirely unsupervised) settings.
游戏。长期以来,人们一直在开发显式手工编码的围棋程序,但 AlphaGo Zero(一个观察棋盘原始状态并下棋的卷积网络)如今已成为该游戏中最强的选手。我预计我们将在其他领域看到类似的结果,例如 DOTA 2 或星际争霸。
Games.Explicitly hand-coded Go playing programs have been developed for a long while, but AlphaGo Zero (a ConvNet that looks at the raw state of the board and plays a move) has now become by far the strongest player of the game. I expect we’re going to see very similar results in other areas, e.g. DOTA 2, or StarCraft.
数据库。人工智能之外更传统的系统也看到了转变的早期迹象。例如,“学习索引结构的案例”用神经网络替换了数据管理系统的核心组件,在速度上比缓存优化的 B 树快高达 70%,同时节省了一个数量级的内存。
Databases. More traditional systems outside of Artificial Intelligence are also seeing early hints of a transition. For instance, “The Case for Learned Index Structures” replaces core components of a data management system with a neural network, outperforming cache-optimized B-Trees by up to 70% in speed while saving an order-of-magnitude in memory.
你会注意到,我上面的许多链接都涉及谷歌的工作。这是因为谷歌目前处于将自身大量代码重写为软件 2.0 代码的前沿。“一个模型统治一切”提供了这可能是什么样子的早期草图,其中各个领域的统计强度被融合成一个对世界的一致理解。
You’ll notice that many of my links above involve work done at Google. This is because Google is currently at the forefront of re-writing large chunks of itself into Software 2.0 code. “One model to rule them all” provides an early sketch of what this might look like, where the statistical strength of the individual domains is amalgamated into one consistent understanding of the world.
为什么我们更倾向于将复杂程序移植到软件 2.0?显然,一个简单的答案是它们在实践中表现更好。然而,还有许多其他便利的理由让我们偏好这一技术栈。让我们看看软件 2.0(例如:卷积网络)相比软件 1.0(例如:生产级 C++ 代码库)的一些优势。软件 2.0 具有以下特点:
Why should we prefer to port complex programs into Software 2.0? Clearly, one easy answer is that they work better in practice. However, there are a lot of other convenient reasons to prefer this stack. Let’s take a look at some of the benefits of Software 2.0 (think: a ConvNet) compared to Software 1.0 (think: a production-level C++ code base). Software 2.0 is:
免费加入 Medium 即可获取此作者的最新动态。
Join Medium for free to get updates from this writer.
计算同质性。典型的神经网络,粗略来说,仅由两种操作堆叠而成:矩阵乘法和零阈值化(ReLU)。相比之下,经典软件的指令集则显著更加异构和复杂。由于只需为少数核心计算原语(如矩阵乘法)提供软件 1.0 实现,因此更容易做出各种正确性/性能保证。
Computationally homogeneous. A typical neural network is, to the first order, made up of a sandwich of only two operations: matrix multiplication and thresholding at zero (ReLU). Compare that with the instruction set of classical software, which is significantly more heterogenous and complex. Because you only have to provide Software 1.0 implementation for a small number of the core computational primitives (e.g. matrix multiply), it is much easier to make various correctness/performance guarantees.
易于集成到芯片中。作为推论,由于神经网络的指令集相对较小,将这些网络实现得更接近硅片(例如使用定制 ASIC、神经形态芯片等)要容易得多。当低功耗智能在我们周围普及时,世界将发生改变。例如,小型、廉价的芯片可以集成预训练的卷积网络、语音识别器和 WaveNet 语音合成网络,形成一个可附着在物品上的小型原型大脑。
Simple to bake into silicon. As a corollary, since the instruction set of a neural network is relatively small, it is significantly easier to implement these networks much closer to silicon, e.g. with custom ASICs, neuromorphic chips, and so on. The world will change when low-powered intelligence becomes pervasive around us. E.g., small, inexpensive chips could come with a pretrained ConvNet, a speech recognizer, and a WaveNet speech synthesis network all integrated in a small protobrain that you can attach to stuff.
恒定运行时间。典型神经网络前向传播的每次迭代所需的 FLOPS 完全相同。与代码可能通过某个庞大的 C++ 代码库采取的不同执行路径相比,这里零可变性。当然,你可以有动态计算图,但执行流程通常仍然受到显著约束。这样,我们也几乎保证永远不会陷入意外的无限循环。
Constant running time. Every iteration of a typical neural net forward pass takes exactly the same amount of FLOPS. There is zero variability based on the different execution paths your code could take through some sprawling C++ code base. Of course, you could have dynamic compute graphs but the execution flow is normally still significantly constrained. This way we are also almost guaranteed to never find ourselves in unintended infinite loops.
恒定内存使用。与上述相关,任何地方都没有动态分配的内存,因此也几乎不可能出现交换到磁盘或需要在代码中查找的内存泄漏。
Constant memory use. Related to the above, there is no dynamically allocated memory anywhere so there is also little possibility of swapping to disk, or memory leaks that you have to hunt down in your code.
高度可移植。与经典二进制文件或脚本相比,矩阵乘法序列在任意计算配置上运行要容易得多。
It is highly portable. A sequence of matrix multiplies is significantly easier to run on arbitrary computational configurations compared to classical binaries or scripts.
非常灵活。如果你有一段 C++ 代码,有人希望你将速度提高一倍(必要时以性能为代价),那么针对新规格调整系统将非常复杂。然而,在软件 2.0 中,我们可以取出网络,移除一半的通道,重新训练,然后——它恰好以两倍的速度运行,效果稍差。这很神奇。相反,如果你碰巧获得更多数据/算力,你可以立即通过增加更多通道并重新训练来使程序工作得更好。
It is very agile. If you had a C++ code and someone wanted you to make it twice as fast (at cost of performance if needed), it would be highly non-trivial to tune the system for the new spec. However, in Software 2.0 we can take our network, remove half of the channels, retrain, and there — it runs exactly at twice the speed and works a bit worse. It’s magic. Conversely, if you happen to get more data/compute, you can immediately make your program work better just by adding more channels and retraining.
模块可以融合成最优整体。我们的软件通常分解为通过公共函数、API 或端点通信的模块。然而,如果两个最初分别训练的软件 2.0 模块交互,我们可以轻松地通过整个系统进行反向传播。想想看,如果你的网络浏览器能够自动重新设计底层系统指令(向下 10 层)以实现更高的网页加载效率,那该有多棒。或者你导入的计算机视觉库(如 OpenCV)可以根据你的特定数据自动调整。在软件 2.0 中,这是默认行为。
Modules can meld into an optimal whole. Our software is often decomposed into modules that communicate through public functions, APIs, or endpoints. However, if two Software 2.0 modules that were originally trained separately interact, we can easily backpropagate through the whole. Think about how amazing it could be if your web browser could automatically re-design the low-level system instructions 10 stacks down to achieve a higher efficiency in loading web pages. Or if the computer vision library (e.g. OpenCV) you imported could be auto-tuned on your specific data. With 2.0, this is the default behavior.
它比你更优秀。最后,也是最重要的,在大量有价值的垂直领域中(目前至少包括任何与图像/视频和声音/语音相关的内容),神经网络是比你我所能想出的任何代码都更好的代码。
It is better than you. Finally, and most importantly, a neural network is a better piece of code than anything you or I can come up with in a large fraction of valuable verticals, which currently at the very least involve anything to do with images/video and sound/speech.
2.0 技术栈也有其自身的缺点。在优化结束时,我们得到的是运行良好的大型网络,但很难解释其工作原理。在许多应用领域,我们将面临选择:要么使用我们理解但准确率只有 90% 的模型,要么使用准确率高达 99% 但我们无法理解的模型。
The 2.0 stack also has some of its own disadvantages. At the end of the optimization we’re left with large networks that work well, but it’s very hard to tell how. Across many applications areas, we’ll be left with a choice of using a 90% accurate model we understand, or 99% accurate model we don’t.
2.0 技术栈可能以不直观且令人尴尬的方式失败,或者更糟的是,它们可能“静默失败”,例如,默默采纳训练数据中的偏见,而当模型规模在大多数情况下轻易达到数百万时,这些偏见极难被正确分析和检查。
The 2.0 stack can fail in unintuitive and embarrassing ways ,or worse, they can “silently fail”, e.g., by silently adopting biases in their training data, which are very difficult to properly analyze and examine when their sizes are easily in the millions in most cases.
最后,我们仍在发现这一技术栈的一些奇特性质。例如,对抗样本和攻击的存在凸显了该技术栈的不直观特性。
Finally, we’re still discovering some of the peculiar properties of this stack. For instance, the existence of adversarial examples and attacks highlights the unintuitive nature of this stack.
软件 1.0 是我们编写的代码。软件 2.0 是由基于评估标准(例如“正确分类这些训练数据”)的优化所编写的代码。很可能任何程序不明确但可以反复评估其性能的场景(例如——你是否正确分类了一些图像?你是否赢得了围棋比赛?)都会经历这种转变,因为优化可以找到比人类编写的代码好得多的代码。
Software 1.0 is code we write. Software 2.0 is code written by the optimization based on an evaluation criterion (such as “classify this training data correctly”). It is likely that any setting where the program is not obvious but one can repeatedly evaluate the performance of it (e.g. — did you classify some images correctly? do you win games of Go?) will be subject to this transition, because the optimization can find much better code than what a human can write.
我们看待趋势的视角很重要。如果你将软件 2.0 视为一种新的、新兴的编程范式,而不是仅仅将神经网络视为机器学习技术类别中一个相当不错的分类器,那么外推就会变得更加明显,并且显然还有更多工作要做。
The lens through which we view trends matters. If you recognize Software 2.0 as a new and emerging programming paradigm instead of simply treating neural networks as a pretty good classifier in the class of machine learning techniques, the extrapolations become more obvious, and it’s clear that there is much more work to do.
特别是,我们已经构建了大量的工具来帮助人类编写 1.0 代码,例如具有语法高亮、调试器、性能分析器、转到定义、git 集成等功能的强大 IDE。在 2.0 栈中,编程是通过积累、处理和清理数据集来完成的。例如,当网络在某些困难或罕见情况下失败时,我们不会通过编写代码来修复这些预测,而是通过包含更多这些情况的标记示例。谁将开发第一个软件 2.0 IDE,帮助完成积累、可视化、清理、标记和采购数据集的所有工作流程?也许 IDE 会基于每个示例的损失弹出网络怀疑标记错误的图像,或者通过用预测种子标记来辅助标记,或者根据网络预测的不确定性建议有用的示例进行标记。
In particular, we’ve built up a vast amount of tooling that assists humans in writing 1.0 code, such as powerful IDEs with features like syntax highlighting, debuggers, profilers, go to def, git integration, etc. In the 2.0 stack, the programming is done by accumulating, massaging and cleaning datasets. For example, when the network fails in some hard or rare cases, we do not fix those predictions by writing code, but by including more labeled examples of those cases. Who is going to develop the first Software 2.0 IDEs, which help with all of the workflows in accumulating, visualizing, cleaning, labeling, and sourcing datasets? Perhaps the IDE bubbles up images that the network suspects are mislabeled based on the per-example loss, or assists in labeling by seeding labels with predictions, or suggests useful examples to label based on the uncertainty of the network’s predictions.
类似地,GitHub 是软件 1.0 代码非常成功的家园。是否存在软件 2.0 GitHub 的空间?在这种情况下,仓库是数据集,提交由标签的添加和编辑组成。
Similarly, Github is a very successful home for Software 1.0 code. Is there space for a Software 2.0 Github? In this case repositories are datasets and commits are made up of additions and edits of the labels.
传统的包管理器和相关的服务基础设施,如 pip、conda、docker 等,帮助我们更轻松地部署和组合二进制文件。我们如何有效地部署、共享、导入和使用软件 2.0 二进制文件?神经网络的 conda 等价物是什么?
Traditional package managers and related serving infrastructure like pip, conda, docker, etc. help us more easily deploy and compose binaries. How do we effectively deploy, share, import and work with Software 2.0 binaries? What is the conda equivalent for neural networks?
短期内,软件 2.0 将在任何可以反复且低成本进行评估、且算法本身难以显式设计的领域中变得越来越普遍。有许多令人兴奋的机会来考虑整个软件开发生态系统以及如何使其适应这种新的编程范式。从长远来看,这种范式的未来是光明的,因为越来越清楚的是,当我们开发 AGI 时,它肯定是用软件 2.0 编写的。
In the short term, Software 2.0 will become increasingly prevalent in any domain where repeated evaluation is possible and cheap, and where the algorithm itself is difficult to design explicitly. There are many exciting opportunities to consider the entire software development ecosystem and how it can be adapted to this new programming paradigm. And in the long run, the future of this paradigm is bright because it is increasingly clear that when we develop AGI, it will certainly be written in Software 2.0.
[](https://karpathy.medium.com/?source=post_page---post_author_info--a64152b37c35---------------------------------------)
[](https://karpathy.medium.com/?source=post_page---post_author_info--a64152b37c35---------------------------------------)
我喜欢在大型数据集上训练深度神经网络。
I like to train deep neural nets on large datasets.